# Canonical book model and rebuilds

The canonical book model is the system's shared record of what the book contains and how each part becomes audio and navigation. “Canonical” means that downstream stages refer to the same structured representation; it does not mean that an AI transcript replaces the original source.

**Proposed invariant: every source passage is represented in the output or has an explicit, reviewable reason for exclusion.** This is the strongest protection against quiet omission, duplication and paraphrasing.

## Follow one passage through the system

<!-- diagram:book-model -->

## What to retain

| Record | Essential information | Why it matters |
|---|---|---|
| Source artifact | Original bytes, hash, format, ingestion version | Reproduce and audit a conversion. |
| Source span | File and stable structural location; original text or audio range | Locate the evidence behind every transformation. |
| Structure node | Heading, paragraph, note, list, table, page marker; parent and order | Preserve reading order and navigation semantics. |
| Text transformation | Original text, approved spoken form, rule, rationale, reviewer | Distinguish faithful normalization from accidental rewriting. |
| Speech recipe | Model/weights version, voice asset, lexicon version, settings | Explain and selectively regenerate the audio. |
| Audio mapping | Asset hash, sample rate, start/end sample positions, source span IDs | Tie content and navigation to actual media. |
| Quality record | Check type, artifact version, evidence, defect status, resolution | Prevent stale approvals from surviving changes. |
| Output build | Profile version, tool versions, dependencies, package hashes | Explain exactly what was delivered. |

This schema is a proposal. Its final field names and storage technology are implementation choices.

## Preserve three forms of text

Store original written text, approved spoken text, and observed ASR text separately. Each answers a different question: what did the source say, what should be spoken, and what did a recognizer hear?

Example:

```text
Original text:  $1,205
Spoken text:    one thousand two hundred five dollars
Transformation: currency/USD
Source location: chapter 4, paragraph 12, characters 18–24
ASR evidence:   populated only after examining the generated audio
```

A normalization rule should never convert “fifty” and “fifteen” into the same comparison token. Preserve high-consequence details such as negation, units, dates and proper names.

## Use stable IDs and immutable versions

Give structural nodes stable identities across edits when they represent the same passage. Give source snapshots and generated assets immutable version IDs or content hashes. A node can point to a new approved audio asset without overwriting the old one.

A timestamp alone is an inadequate identity: earlier edits can shift every later timestamp. Store the logical destination and the resolved audio position, then derive the latter for each completed build.

## Track two independent audio references

A navigation item points to a destination in the main narration and may also reference a spoken label in a separate headings file. Record both asset identities and timing ranges explicitly. Confusing the two timelines can produce a correct-sounding label that jumps to the wrong place.

## Explore which work changes

<!-- diagram:repair -->

## Worked repair example

1. An operator corrects a surname in the book lexicon.
2. Dependency lookup finds the passages whose spoken forms use that entry.
3. Those passages receive new speech recipes and audio versions.
4. The affected assembled masters and their final sample positions change.
5. Dependent SMIL/NCX timing, durations, checksums, protection and packages are rebuilt.
6. Validation runs against the new artifact hashes; the prior build remains auditable.

The system should record who accepted the correction, why it was made, and whether all occurrences should change. A name may intentionally have different pronunciations in different contexts.

## Define cache identity carefully

Include source/spoken text, model and voice artifacts, normalization and lexicon versions, generation parameters, and relevant processing versions. Store the accepted audio itself. A random seed cannot guarantee byte-identical future synthesis across different hardware or software versions.
