# System architecture

The proposed system has two front ends and one production core. Electronic text needs parsing and narration. Existing audiobooks need structure recovery and supplemental speech. Both converge on the same book representation, audio processing, compiler, validation, and delivery stages.

## How the parts connect

<!-- diagram:architecture -->

## Who does what

| Work | Proposed approach | Reason |
|---|---|---|
| Recognize uncertain structure and ambiguous text | Deterministic parsing first, model-assisted proposals when useful | Source evidence stays available and uncertainty can be reviewed. |
| Generate narration and announcements | Approved, replaceable speech models | Voices and model families can be evaluated without rewriting the compiler. |
| Find possible spoken defects | Independent ASR, alignment, acoustic checks, targeted review | No single recognizer is a perfect judge. |
| Maintain IDs, source coverage, references and timing | Typed data structures and conventional code | These properties need exact, repeatable checks. |
| Construct XML, encode, protect and package | Versioned deterministic compiler plus approved media tools | Compatibility depends on exact rules and artifacts. |
| Accept a book | Agreed validators, playback tests and authorized review | Passing a model check is only one source of evidence. |

## The main modules

**Ingestion** stores the original files and parses format-specific structure. **The book service** manages immutable versions, source spans, metadata claims and approvals. **Speech and analysis workers** synthesize or inspect bounded segments. **The compiler** uses accepted content and a frozen timeline to emit the required files. **Independent validators** report defects. **The review interface** lets an operator resolve exceptions and trigger dependent rebuilds.

The **delivery adapter** publishes only accepted, complete outputs through the agreed NLS interface. It should not infer an interface from the fact that the environment uses AWS.

## Separate three kinds of segmentation

A speech chunk is a unit that can be regenerated. A navigation element is a destination a reader can select. A delivered audio file is a packaging and playback unit. They do not need to have matching boundaries.

For illustration, one chapter could contain 80 synthesis chunks, 12 navigable sections, and one delivered chapter audio file. Those numbers are explanatory, not requirements. Keeping the concepts separate enables local repairs while respecting file limits and meaningful navigation.

<!-- diagram:segmentation -->

## A book behaves like an incremental build

A pronunciation correction can affect several speech chunks. Regenerate those chunks, reassemble the affected masters, recalculate the dependent timeline, rebuild navigation/XML, encode and protect the changed content, then update checksums and packaging. Unaffected accepted synthesis can stay cached.

This is a dependency graph: each output records the exact inputs from which it was made. If an input changes, downstream evidence becomes stale. The build system should never keep a green validation result for an old version of a file.

## Boundaries for AI assistance

Model outputs are proposals or generated media until verified. They must not silently delete source text, invent page numbers, substitute a narrator, write arbitrary final XML, or mark their own output accepted. Books and metadata are untrusted content, including instructions that appear inside them.

A useful model response is structured: a proposed chapter boundary, the source span supporting it, alternative interpretations, and confidence evidence. The compiler consumes an accepted decision, not a free-form conversation.

## Small services before many services

The module boundaries can initially live in one application with separate worker processes. Splitting every module into a network service is not necessary to prove the approach. Establish stable data contracts, job IDs and reproducible artifacts first; deploy independently where isolation or scale justifies it.
