**Yes—I consider this technically feasible with current AI models and existing software components.** The attachments point to a specific architecture: **an AI-assisted audiobook production system with a deterministic compiler at its center.** AI should recover structure, interpret difficult text, generate narration, and identify likely errors. Conventional code should control source fidelity, timing, navigation, standards compliance, encryption, and packaging. The most promising strategy is not to make one powerful agent “produce a talking book.” It is to build a system that can **generate, verify, repair, and rebuild individual portions of a book without losing track of the source or invalidating everything else.** I reviewed the draft SOW and J1–J3 and researched the relevant tools and current models. This is an architecture and feasibility assessment; I have not yet benchmarked the models against these requirements or demonstrated a package on an NLS player. ## 1. What the attachments establish—and where they conflict ### There are two production paths, but they should share one core The SOW requires both: | Input | What the system must do | |---|---| | **Existing commercial audiobook** | Preserve the supplied narration, identify navigable divisions, generate supplemental announcements and navigation audio, and construct the DTB. | | **Electronic text** | Recover the book’s structure, synthesize the narration, and construct the DTB. | Both require metadata, navigation, WAV/AMR-WB+ processing as applicable, protection, validation, and delivery. The text inputs explicitly include EPUB, Word, RTF, and plain text. Scanned-book OCR is not an explicit core requirement in this draft. SOW-TTS-DTB_draft.docx[Draft SOW, §§3.1–3.2](sandbox:/mnt/data/nls_attachments/SOW-TTS-DTB_draft.docx) **Engineering implication:** build two ingestion/front-end workflows feeding the same internal book model, audio processing system, and output compiler—not two independent applications. ### The standards versions need to become explicit configuration There is a real discrepancy in the package: - The SOW repeatedly specifies **ANSI/NISO Z39.86-2002**. - J1 is **NLS Specification 1203:2018**, version 1.3, January 2019. - J2 is actually **ANSI/NISO Z39.86-2005, reaffirmed 2012**. - J3 is the December 2018 navigation guidance. J1 Digital Talking Book Construction v1.3.pdf[J1, cover](sandbox:/mnt/data/nls_attachments/J1%20Digital%20Talking%20Book%20Construction%20v1.3.pdf) · J2 ANSI-NISO Digital Talking Book Specifications.pdf[J2, cover](sandbox:/mnt/data/nls_attachments/J2%20ANSI-NISO%20Digital%20Talking%20Book%20Specifications.pdf) · J3 Navigation Markers Instructions.pdf[J3, cover](sandbox:/mnt/data/nls_attachments/J3%20Navigation%20Markers%20Instructions.pdf) This is not just a date discrepancy. For example, the 2005 standard uses `pageList`/`pageTarget` for page navigation; the 2002 format uses the older navigation representation. A generic “DAISY 3 exporter” is therefore not automatically the required NLS exporter. [The DAISY Consortium](https://www.daisy.org/z3986/specifications/Z39-86-2002.html) The current public NLS construction specification is **1203:2025**, and it still references Z39.86-2002. It also changes some older constraints, including removing the 250-file limit. Meanwhile, the SOW calls for “current applicable versions.” [The Library of Congress](https://www.loc.gov/nls/who-we-are/guidelines-and-specifications/contract-specifications/digital-talking-book-requirements/) **My approach:** define a versioned **NLS output profile** containing the approved schemas, metadata rules, navigation rules, audio parameters, filenames, and packaging requirements. Every generated book records the exact profile used. Do not silently mix requirements from different editions. One additional ambiguity: the SOW requests “various multimedia types,” while J1’s metadata requirements prescribe `audioNCX`. I would implement the NLS audio-plus-navigation profile first, while retaining enough internal text structure to support an explicitly requested full-text-and-audio profile later. J1 Digital Talking Book Construction v1.3.pdf[SOW, §3.2; J1, §5.3](sandbox:/mnt/data/nls_attachments/J1%20Digital%20Talking%20Book%20Construction%20v1.3.pdf) ### Self-hosted inference is the natural default The draft requires a non-subscription solution without recurring software/service access requirements and with Library ownership of the delivered software. SOW-TTS-DTB_draft.docx[SOW, §§2 and 3.12](sandbox:/mnt/data/nls_attachments/SOW-TTS-DTB_draft.docx) My technical interpretation is to **avoid a mandatory commercial TTS API in the production architecture**. Deliver the application, model artifacts, configuration, and deployment machinery into the Library’s environment. That does not automatically resolve rights to third-party model weights, libraries, or voices. Those need an approved dependency inventory. But it strongly favors an architecture that continues functioning without contacting the contractor or a model vendor. ## 2. The architecture I would build ```text Electronic text Commercial audiobook │ │ Document parsing Audio + metadata ingestion Structure recovery Transcription / alignment Text normalization Structure recovery │ │ └──────────── Canonical book model ────────┘ │ Narration and announcement plan │ Speech generation where needed │ Verification and local repair │ Finalized audio timeline │ Deterministic NLS DTB compiler │ Encoding → protection → packaging │ Independent validation and playback QA │ NLS delivery interface ``` ### The central data model is the most important engineering decision I would maintain a structured representation of the book containing: | Area | Information retained | |---|---| | **Source and provenance** | Original files, hashes, source locations, catalog metadata, and conflicting metadata claims. | | **Book structure** | Chapters, headings, paragraphs, notes, references, lists, tables, and authoritative page markers. | | **Text** | Original text, approved spoken form, generated announcements, and explanations for transformations. | | **Speech configuration** | Model version, approved voice, pronunciation lexicon, and generation parameters. | | **Audio mapping** | Which source spans correspond to which audio assets and sample ranges. | | **Quality state** | Validation results, detected discrepancies, confidence evidence, repairs, and approvals. | | **Output configuration** | The specific NLS/DAISY profile and build version. | The governing invariant should be: > **Every source passage is either represented in the output or has an explicit, reviewable reason for exclusion.** An LLM should never quietly summarize a paragraph, “clean up” unfamiliar wording, or remove text that it considers unimportant. ### Keep three different granularities separate **Speech-generation chunks, navigation elements, and delivered audio files should not be the same thing.** For example, a chapter might contain 80 synthesis chunks, 12 navigable sections, and one delivered chapter audio file. Internal sentence-level alignment is useful for checking and repairing speech, but it does not mean publishing a navigation marker or SMIL element for every word. That separation is what makes the system both repairable and compatible with the delivery constraints. ### Treat a book as an incremental build Suppose an operator corrects the pronunciation of a name. The system should identify all affected speech chunks, regenerate them, rebuild the affected audio masters, recalculate dependent timestamps and checksums, and repackage the book. Unaffected synthesis should remain cached. That is much closer to a compiler/build system than a chat-agent workflow. It also gives AI-assisted implementation a strong foundation: individual transformations have explicit inputs, outputs, and testable invariants. ## 3. Recovering structure and preparing the source ### Electronic text: preserve existing structure before asking AI to infer it I would use format-specific importers: **EPUB:** parse the package, reading order, XHTML body, headings, links, notes, and any authoritative page mapping. **Word:** inspect heading styles, outline levels, numbered lists, tables, and footnote/endnote relationships. Do not reduce the document to plain text first. **RTF and legacy Word formats:** normalize through a sandboxed conversion step, while retaining the original and recording what structural information survived. **Plain text:** apply deterministic heuristics and then an LLM to propose headings, lists, or special passages. Record the evidence and uncertainty because some structure has genuinely been lost. The supplied navigation instructions explicitly say that the **main content—not merely the table of contents—must determine navigation**. They also distinguish hierarchical navigation from other structures such as notes and segments. J3 Navigation Markers Instructions.pdf[J3, §§3.1–3.2](sandbox:/mnt/data/nls_attachments/J3%20Navigation%20Markers%20Instructions.pdf) That is a good use of modern models: classifying source passages and resolving ambiguous structure. But each proposed division should point back to a specific source location. No invented page numbers or headings. ### Text normalization needs its own subsystem The spoken representation of: > “Dr. Smith paid $1,205 on 03/04/2026.” requires several decisions. What does “Dr.” mean here? How should the amount be spoken? Which date convention applies? I would combine a deterministic normalizer with contextual disambiguation. NVIDIA’s NeMo text-processing project provides a reusable written-to-spoken normalization foundation. [GitHub](https://github.com/NVIDIA/NeMo-text-processing?utm_source=chatgpt.com) The proposed workflow would preserve both forms: ```text Source: "$1,205" Spoken form: "one thousand two hundred five dollars" Rule: currency/USD Source span: chapter-04, paragraph-12, characters 18–24 ``` For names and specialized vocabulary, maintain a **book-level pronunciation lexicon**. Corrections should propagate consistently throughout the book. Do not assume every neural TTS model accepts the same SSML or phoneme notation. Each engine adapter should explicitly expose the controls it actually supports. It should also escape or neutralize model-specific control tokens in ordinary source text so that a literal bracketed phrase is not accidentally interpreted as a performance instruction. ### Commercial audio: use evidence in descending order of reliability For existing recordings, I would combine: **Publisher metadata and chapter boundaries → matching electronic text, when available → ASR transcript → semantic structure classification → precise audio alignment.** An LLM can help distinguish a genuine chapter announcement from a character saying “we discussed that in chapter three.” But it should not decide the final timestamp by itself. Qwen3-ASR provides both transcription models and a separate forced aligner. Its documented aligner window is limited, so long books would need bounded, overlapping alignment jobs rather than one whole-book call. WhisperX is another useful alignment implementation and an independent alternative. [GitHub](https://github.com/QwenLM/Qwen3-ASR) The commercial path should **preserve the supplied narration**. Its transcript is an analysis aid, not an authoritative replacement for the recording. Only announcements, labels, and other required supplemental speech should be synthesized. There is one information limitation worth stating plainly: a recording without accompanying text cannot reliably reveal a printed heading that was never spoken or an original print page number that was never announced. The system should represent that as unavailable information, not have AI manufacture it. ## 4. Speech generation: models, selection, and long-form reliability ### My initial model evaluation set I would build a replaceable TTS interface and test these families against the same representative material: | Candidate | Verified characteristics | Role in the evaluation | |---|---|---| | **Qwen3-TTS** | 0.6B/1.7B variants, local inference, multiple voice-generation modes; the examined 1.7B Base weights are Apache-2.0. | Main candidate for controllable narration and approved custom voices. | | **Chatterbox** | Turbo is a 350M English model; Multilingual V3 is 500M. The official project is MIT-licensed. | Alternative architecture for narration, voice consistency, and difficult passages. | | **Kokoro-82M** | Small, open-weight model with Apache-licensed weights. | Lightweight baseline for both narration and announcement generation. | These are candidates to measure, not a claim that one already satisfies NLS quality requirements. [Hugging Face](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base) I also investigated **Fish Audio S2 Pro**. It is technically relevant, but its model card requires a separate license for commercial use. I would not make it a default dependency without resolving those terms against the SOW. [Hugging Face](https://huggingface.co/fishaudio/s2-pro) My starting implementation would use **Qwen3-TTS as the first adapter, Chatterbox as an independent challenger, and Kokoro as a compact baseline**. The final choice should be based on whole-book behavior and correction effort—not promotional demos or time-to-first-audio. ### How I would generate a full book I would not begin by generating an entire book in one model call, even with a model capable of long outputs. Instead, establish a book-level narration plan: approved narrator voice, pronunciation decisions, language handling, and restrained pacing/style settings. Generate paragraph- or sentence-group chunks, with chunk size tuned experimentally. A reasonable initial experiment would compare chunks representing roughly **20–60 seconds of speech**, keeping sentence and paragraph boundaries intact. That is a proposed tuning range, not a requirement or a claimed optimum. The important tests are whether the model maintains voice and pacing across chunks, handles context correctly, and produces clean joins. Very short chunks may sound disconnected; very long chunks make correction more expensive. The evaluation should find the useful middle ground. **Default narration should prioritize faithful reading over creative performance.** Multiple selectable voices do not imply automatic character acting, added laughs, dramatic paraphrasing, or invented sound effects. ### Use bounded generate–verify–repair loops After synthesizing a chunk: 1. Transcribe it with an independent recognizer and compare it with the approved spoken text. 2. Check alignment and acoustic properties. 3. Regenerate or adjust the chunk when a specific discrepancy is found. 4. Escalate unresolved cases after a bounded number of attempts. A model should not silently switch to a different narrator halfway through a book to resolve a difficult sentence. Fallback policies need to preserve voice consistency or trigger a broader, deliberate regeneration. Approved voice assets also need provenance and permission. I would not automatically clone the narrator of an incoming commercial audiobook merely because the recording is available. ## 5. The deterministic compiler: where compliance becomes concrete This is the part I would specify most rigorously. ### Audio mastering and the final timeline Use integer sample positions internally. Apply approved resampling, loudness processing, padding, and assembly before freezing the final timeline used by the compiler. The supplied J1 requires navigation clips to begin **80–120 milliseconds before narration** and end **150–300 milliseconds afterward**. Its timing model references the source WAV and discusses compensation for compressed-audio offsets. J1 Digital Talking Book Construction v1.3.pdf[J1, §§3.2–3.4](sandbox:/mnt/data/nls_attachments/J1%20Digital%20Talking%20Book%20Construction%20v1.3.pdf) That means coarse ASR timestamps are insufficient. I would combine forced alignment with local waveform/voice-activity analysis to locate the actual speech boundary. A proposed default would place clip boundaries near the middle of the permitted windows, with validation against the finalized audio. Where the required silent margin does not exist, the system must either make an approved audio edit and recalculate dependent timing or flag the case. It should not emit a knowingly invalid timestamp. There is also a TTS-specific clarification to obtain. The current mastering specification requires 44.1 kHz PCM masters and includes requirements about initial recording format; it also specifies integrated loudness of **−22 LUFS ±2** and says heading clips are copied from primary audio. A neural model’s native output and the SOW’s synthesized navigation labels may require an agreed interpretation. Simply resampling does not prove compliance with an “initially recorded” requirement. [The Library of Congress](https://www.loc.gov/nls/who-we-are/guidelines-and-specifications/contract-specifications/audiobook-mastering/) ### Separate the main narration timeline from the headings timeline Navigation labels may use a separate headings audio file. A timestamp in that file is not a timestamp in the main narration. The internal representation should therefore distinguish, for each navigation item: ```text Navigation destination → main narration asset and sample position Spoken navigation label → headings asset and sample range ``` This avoids an easy class of errors: a chapter label that sounds correct but navigates to the wrong place, or a label whose timestamps accidentally reference the wrong audio asset. ### Generate XML from typed structures, not model output The compiler should serialize OPF, NCX, SMIL, and any required supporting files from the canonical representation. Validation should cover both syntax and semantics: identifiers, references, ordering, navigation classes, clip ranges, metadata consistency, required announcements, and correspondence with actual media files. Several details in J1 deserve dedicated tests: | Requirement | Implementation consequence | |---|---| | SMIL file-size and count limits | Partition based on actual serialized bytes, not an arbitrary chapter count. | | Notes and segments have distinct navigation behavior | Model them explicitly rather than turning everything into chapters. | | Separate audio navigation labels | Maintain separate label media and timestamps. | | The `.md5` file is an **XML checksum document** | Do not substitute ordinary `md5sum` output. | | DTDs/entities accompany the package | Bundle approved schema dependencies and resolve them locally. | | Duration metadata follows the specified playback structure | Calculate from the compiled timeline, not a casual sum of file lengths. | These are implementation rules, not AI judgments. J1 Digital Talking Book Construction v1.3.pdf[J1, §§3.3–3.10 and §5.3](sandbox:/mnt/data/nls_attachments/J1%20Digital%20Talking%20Book%20Construction%20v1.3.pdf) For XML processing, I would allow only the packaged, approved DTD/entity dependencies—not arbitrary external entity resolution. ### Reuse DAISY Pipeline selectively DAISY Pipeline already provides document conversion, validation, and speech-enabled workflows. That is useful infrastructure to evaluate rather than recreate indiscriminately. [Daisy](https://daisy.github.io/pipeline/Get-Help/User-Guide/Scripts/) But its DTBook-to-DAISY3 script documents a **2005-3 DTBook input**. I would not assume it is a drop-in NLS 2002 production compiler. [Daisy](https://daisy.github.io/pipeline/Get-Help/User-Guide/Scripts/dtbook-to-daisy3/) My preference would be to reuse suitable importers and transformations while retaining explicit ownership of the **NLS profile, export logic, and acceptance tests**. ### AMR-WB+ and protection should be proved early The required codec is **AMR-WB+**, not ordinary AMR-WB. The NLS profile also constrains the 3GP container and its metadata. This is more specific than “convert the WAV to a compressed audio file.” J1 Digital Talking Book Construction v1.3.pdf[J1, §3.2](sandbox:/mnt/data/nls_attachments/J1%20Digital%20Talking%20Book%20Construction%20v1.3.pdf) I found useful evidence that this need not start from scratch: Hindenburg’s documentation describes an **NLS-supplied encoder component**, and APH’s Book Wizard documentation describes integration with an NLS protector and validators. These establish a precedent for reusable components, although they do not establish what NLS currently supplies or supports. [hindenburg.uservoice.com](https://hindenburg.uservoice.com/knowledgebase/articles/1851514-narrator-nls-export-problem-could-not-locate-a) I would request the current encoder, protector, validators, operating-system requirements, and permitted deployment model before choosing an implementation. The fallback is to evaluate an appropriately licensed encoder implementation and build the required container/protection adapters. AI-assisted coding is useful there, but interoperability still has to be demonstrated against known-good files and actual playback systems. **PDTB protection is separate from cloud storage encryption.** The DAISY protection scheme involves protected content and authorization/key objects; applying an S3 encryption setting does not produce a protected DTB. [The DAISY Consortium](https://daisy.org/activities/standards/pdtb/daisy-protected-digital-talking-book-specification/) ### Packaging has legacy details that deserve exact tests The public NLS delivery specification describes an outer package containing protected and unprotected DTB archives plus WAV masters. It requires uncompressed ZIP entries and particular ZIP-header behavior, including known CRCs and sizes rather than the streaming data-descriptor mechanism. [The Library of Congress](https://www.loc.gov/nls/who-we-are/guidelines-and-specifications/contract-specifications/dtb-delivery-requirements/) I would make packaging a tested module, write to a seekable staging artifact, validate it, and only then publish it to S3. Logs and QA reports should remain separate unless the final interface expressly permits them inside the deliverable. ## 6. Quality assurance should be the automation engine The architecture should be designed around **detectable defects and precise repairs**, not merely producing an output file. ### Separate the kinds of correctness | Layer | What it verifies | |---|---| | **Source fidelity** | No unexplained missing, duplicated, reordered, or altered content. | | **Speech fidelity** | Generated speech corresponds to the approved spoken text. | | **Acoustic quality** | Loudness, clipping, silence, joins, intelligibility, and consistency. | | **Navigation correctness** | Labels and destinations correspond to actual source divisions. | | **Format correctness** | Schemas, cross-references, timing, encoding, checksums, and packaging. | | **Playback correctness** | The protected output behaves correctly on intended NLS players/apps. | These layers should fail independently and produce specific diagnostics. ### ASR is useful, but not a proof of perfect narration An important distinction: **Forced alignment answers approximately “where does this expected text fit in the audio?” It does not independently prove that the model actually spoke every expected word.** Likewise, a transcription mismatch can be an ASR error rather than a TTS error. WhisperX itself documents alignment limitations, including words that cannot be aligned by the selected alignment model. [GitHub](https://github.com/m-bain/whisperX) My proposed checker would combine independent transcription, alignment, acoustic evidence, and protected-token checks for names, quantities, dates, and negation. Normalization must not erase meaningful differences—for example, treating “fifteen” and “fifty” as interchangeable. For high-risk discrepancies, an additional recognizer or audio-capable model can provide another opinion. But agreement among models should be treated as evidence, not an absolute guarantee. ### The review interface should expose exceptions, not force a complete restart For each flagged issue, show the source passage, expected spoken form, actual transcript, a short audio excerpt, waveform/timing information, and the reason for the flag. Useful actions would include correcting pronunciation, changing a boundary, accepting a documented exception, or regenerating a bounded passage. Each action should trigger the necessary dependent rebuilds. The metrics I would use to judge automation are **unresolved content errors, navigation accuracy, correction attempts, and review minutes per accepted audio hour**, segmented by input type. A beautiful sample is not enough; evaluate complete books and the difficult transitions within them. Accessibility testing should also exercise chapter navigation, note skipping, resume behavior, and faster playback—not just listen to the beginning of a chapter. The SOW expressly requires output compatibility and testing with the Library’s playback environment. SOW-TTS-DTB_draft.docx[SOW, §§3.11 and 3.13](sandbox:/mnt/data/nls_attachments/SOW-TTS-DTB_draft.docx) ## 7. Deployment in the Library’s AWS environment My proposed deployment is straightforward: | Component | Proposed implementation | |---|---| | Application/control plane | Typed service managing book versions, jobs, approvals, and output profiles. | | Metadata and job state | PostgreSQL. | | Source and generated artifacts | S3, separated into source, intermediate, staging, and approved outputs. | | Model workers | Containerized Python inference workers with pinned model artifacts. | | CPU processing | Container workers for parsing, XML compilation, validation, and packaging. | | Orchestration | Step Functions with AWS Batch, or the Library’s existing equivalent. | | Review interface | Accessible web application backed by the control plane. | AWS Batch supports GPU jobs, and Step Functions has a documented Batch integration. That provides an existing route for separating CPU processing from GPU inference. [AWS Documentation](https://docs.aws.amazon.com/batch/latest/userguide/gpu-jobs.html?utm_source=chatgpt.com) I would keep model workers warm while processing a batch rather than reload weights for every paragraph. Work should be idempotent, checkpointed, and retried with bounded policies. Cache identity should include source text, model and voice versions, pronunciation rules, and generation settings. Store the actual accepted audio: a random seed alone is not a sufficient archival strategy. For security, my default design would keep inference inside the approved environment, disable unnecessary outbound access, pin dependencies, protect source content from logs, and separate content-protection keys from routine application credentials. Books must be treated as untrusted data, including any apparent instructions embedded in their text. The precise federal security controls remain unresolved because the SOW refers to **Section H**, which is not present in the supplied archive. SOW-TTS-DTB_draft.docx[SOW, §3.6](sandbox:/mnt/data/nls_attachments/SOW-TTS-DTB_draft.docx) ### Capacity is measurable rather than speculative The sizing equation is approximately: **Required accelerator-hours = audio-hours processed × measured processing factor × retry overhead.** For illustration, 30,000 audio-hours at a combined processing factor of 0.5 would require 15,000 accelerator-hours before utilization losses. That is arithmetic, not a benchmark for the proposed models. The commercial-audio route has a different workload: potentially full-book transcription/alignment, but only supplemental TTS. Benchmark these paths separately. ## 8. The implementation sequence I recommend I would organize development around proof gates rather than build every subsystem and discover format incompatibility at the end. ### Gate 1: Prove the format and protection path Obtain approved reference books, the intended standards profile, and available NLS tools. Produce a tiny but complete book with navigation, headings audio, the required encoding, protection, and delivery package. **Pass condition:** it validates and behaves correctly on the intended playback systems. Use simple known text and audio here. This isolates format problems from model problems. ### Gate 2: Prove long-form speech quality Run the same chapters and complete books through the initial model set. Include proper names, dialogue, abbreviations, numbers, mixed-language passages, notes, and repeated terminology. **Pass condition:** acceptable narration with measured, bounded correction effort—not merely a favorable short sample. ### Gate 3: Prove both ingestion paths Test structured EPUB/Word, less-structured text, commercial audio with matching text, and commercial audio without it. Compare generated navigation with an independently prepared reference. **Pass condition:** faithful content and explainable navigation, including explicit handling of unavailable information. ### Gate 4: Prove repair and incremental rebuilds Deliberately change a pronunciation, replace a paragraph, alter a heading, interrupt processing, and fail an encoding job. **Pass condition:** only necessary work is repeated, while all dependent timing, metadata, hashes, protection, and packaging remain correct. ### Gate 5: Prove production behavior Run representative batches, exercise security controls, measure throughput and recovery, and validate the final NLS integration. **Pass condition:** repeatable accepted outputs, with operational metrics that distinguish processing failures from content exceptions. AI coding agents fit particularly well here: importer implementation, schema-derived tests, malformed-input generation, reference-file comparison, and regression analysis are bounded tasks. I would still keep acceptance tests and reference outputs independent of the agent that wrote the implementation. ## The technical questions I would resolve with NLS Before fixing the design, I would obtain answers to six specific questions: 1. **Which standards editions and multimedia profiles govern?** In particular, the 2002/2005 mismatch and `audioNCX` versus other output types. 2. **Which encoder, protector, validators, keys, and test players will NLS supply?** Include licensing, operating systems, and batch/cloud deployment support. 3. **What accompanies commercial recordings?** Chapter metadata, matching electronic text, and authoritative page references materially change navigation recovery. 4. **How do the mastering rules apply to synthesized audio?** Clarify native sample rates, heading generation versus copying, and synthetic narrator metadata. 5. **What are the exact security and dependency constraints?** Supply Section H and requirements for models, voice assets, network access, and retention. 6. **What constitutes an accepted book?** Define quality thresholds, supported content classes, turnaround targets, and the actual delivery interface. ## Bottom line **I would pursue this as a standards-driven, AI-assisted production pipeline—not as a research project to invent a new speech model.** The concrete starting architecture is: **structured source importers, a provenance-preserving book model, interchangeable self-hosted speech and alignment models, bounded verification/repair loops, a custom NLS-aware compiler, and validated encoding/protection components running in the Library’s AWS environment.** The first decisive demonstration should combine **a correctly protected, navigable book on an NLS player** with **a long-form narration test showing faithful content and manageable automated repair**. Those two demonstrations would resolve far more of the project’s technical uncertainty than a polished interface or a convincing 30-second voice sample.