# Source ingestion and navigation recovery

Ingestion should retain source evidence before attempting to improve it. The workflow differs by input format, but every importer should produce source-linked content and proposed navigation for the common book model.

## Electronic text formats

| Format | Recover first | Watch for |
|---|---|---|
| EPUB | Package metadata, reading order, XHTML structure, headings, links, notes, authoritative page mapping | Visual styling that disguises semantic structure; broken reading order. |
| Word | Heading styles, outline levels, lists, tables, footnotes and endnotes | Direct formatting that looks like a heading but has no heading style. |
| RTF or legacy Word | Convert in an isolated process while keeping the original and a conversion record | Lost notes, tables, symbols, or outline information. |
| Plain text | Delimiters and deterministic heuristics, followed by model-assisted classification | Missing hierarchy and ambiguous headings; uncertain reading order. |

The format list comes from the [SOW §§3.1–3.2](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/34eacfe92b1a46228b9f90674c4c60e1/download); the importer design is proposed. Do not flatten a structured source into plain text before extracting its relationships.

## Main content determines navigation

[J3 §3.1](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/48b52294c2f44f0780f220a9eff56cbd/download) directs navigation work to the main content rather than simply copying the table of contents. A table of contents may omit subheadings or describe them differently. Inspect the actual text and preserve evidence for every proposed division.

J3 also separates hierarchical navigation from nonhierarchical structures such as notes and segments. A note is not necessarily another chapter. The correct distinction affects playback behavior and must survive ingestion into the compiler.

A model can propose that a passage is a section heading, but the proposal should include its source location and supporting evidence. Printed page numbers should come from an authoritative page map or source, not model invention.

## Commercial recordings

Use evidence in this proposed order: publisher metadata and chapter boundaries, matching electronic text when available, an ASR transcript, semantic classification, and precise audio alignment. Record conflicts rather than silently choosing whichever source arrived last.

For example, a model can distinguish a spoken chapter announcement from dialogue mentioning “chapter three.” The final destination still needs alignment and a waveform-level boundary check.

Keep the supplied narration through approved mastering and conversion. The ASR transcript is an analysis aid that can contain errors. Synthesize only the additional speech the approved output requires. Availability of a narrator's recording is not permission to clone the narrator's voice.

A recording cannot reliably reveal an unspoken print heading or an unannounced print page number. Mark these as unavailable, and agree how that input class will be accepted.

## Written text to spoken text

Normalization expands abbreviations, dates, numbers and symbols while retaining original source spans. The sentence “Dr. Smith paid $1,205 on 03/04/2026” needs context: “Dr.” may require disambiguation, the currency needs expansion, and the date convention is ambiguous.

Use deterministic rules for well-defined cases and a reviewable proposal for contextual cases. A book-level pronunciation lexicon keeps repeated terms consistent. [NVIDIA NeMo text processing](https://github.com/NVIDIA/NeMo-text-processing) is a candidate foundation for text normalization, not a guarantee that the whole book is correctly interpreted.

## Model adapters and source safety

Each speech adapter should declare which controls it supports: SSML, phonemes, pronunciation dictionaries, language selection, reference audio, and style instructions. Unsupported controls must fail clearly or use an agreed conversion; they should not be silently ignored.

Treat model-specific control strings in source text as content. Escape or otherwise neutralize them before a model can interpret them as a voice or performance instruction. Retain the literal source and record the transformation.

## What an importer should report

Proposed outputs include an ordered content tree, source coverage map, metadata conflicts, lost-format warnings, navigation proposals and unresolved passages. A successful file parse alone is insufficient: compare recovered content and reading order with the source before moving into speech production.
