This wiki explains a proposed system for the Library of Congress National Library Service for the Blind and Print Disabled (NLS). It would turn electronic books and existing commercial audiobooks into accessible digital talking books (DTBs) with reliable navigation, required audio formats, protection, and delivery packaging.
The central recommendation is an AI-assisted production pipeline with a deterministic compiler at its center. AI proposes structure, prepares speech, and helps detect defects. Conventional software enforces the approved rules and creates the final files. The original analysis considers this feasible; no model benchmark, compliant prototype, or successful NLS-player demonstration is established by the materials here.
Source requirement means a statement found in the supplied draft SOW or a named specification. Edition and applicability still matter. Proposal means an engineering choice from the analysis or this wiki. Open decision means an answer is needed from NLS or from testing before the design can be fixed.
Public specifications are useful comparison material. A newer public edition does not by itself amend this opportunity. The standards page explains the 2002/2005 mismatch and the difference between the attached and current public construction specifications.
What to prove first
Produce a tiny, complete, protected book that passes the approved validators and plays correctly on the intended NLS systems. In parallel, test sustained narration on representative complete books and measure correction effort. These demonstrations address format interoperability and speech quality separately.
About this local edition
Prepared on October 1, 2026, from the supplied analysis, the four opportunity attachments, and linked primary sources. Added explanations include a glossary, explicit acceptance evidence, dependency tracking, milestones, and decision prompts. Nothing in this wiki records answers from NLS or actual performance results.
Wiki content works offline. Attachment citations open the original files on SAM.gov and require an internet connection. The source library also offers local backup copies. Editable articles are stored alongside it as Markdown. Keep the whole folder together when moving it. See the source library for provenance and the glossary for unfamiliar terms.
The draft statement of work describes software for producing NLS digital talking books through two input routes. The engineering proposal is to share the downstream production system between those routes.
The draft background describes approximately 3,000 DTBs produced annually in the existing operation. That is context, not a numerical throughput acceptance target for the new system. SOW §1.
The two required input routes
Route
Core work
Important boundary
Existing commercial audiobook
Ingest audio and metadata, identify divisions, create required supplemental speech, and construct the DTB.
Proposed handling: preserve the supplied narration through approved mastering/conversion. A transcript supports analysis rather than replacing the recording.
Electronic text
Ingest EPUB, Word, RTF, or plain text; recover structure; generate speech; construct the DTB.
Preserve the source and document any change to its spoken form.
The commercial route explicitly accepts WAV, MP3, OPUS and FLAC, and requires AMR-WB+ conversion with MP3 encoding optional. The text route requires WAV and AMR-WB+ outputs, with MP3 encoding optional. These routes and formats come from SOW §§3.1–3.2. Scanned-book OCR is not an explicit core input in the draft. Treat it as a potential extension that needs scope agreement. A raw scan also creates recognition errors that the current text-first workflow would otherwise avoid.
Requirements and proposed evidence
The first two columns summarize the draft; the last is a proposed way to demonstrate completion, not a Government-approved acceptance checklist.
Area
Draft requirement or concern
Proposed evidence
Outputs
NLS-compatible DTBs, navigation, metadata, required audio and related files; §§3.1–3.2
A reference book from each input route, validator results, and playback observations.
Processing
Batch operation and automation with useful operator control
Repeatable batch run, failed-job recovery, and an exception-review demonstration.
Integration
Work within the Library's AWS environment and interfaces
Deployment in the approved environment and a traced delivery to the agreed destination.
Security
Federal security requirements referenced through Section H; §3.6
An agreed controls matrix after the missing section is supplied.
Compatibility and testing
Library playback environment; §§3.11, 3.13
Test records for the agreed hardware, applications, and protected packages.
Ownership and continuity
Non-subscription delivery and Library ownership; §§2, 3.12
Deliverable inventory, reproducible deployment, and dependency/voice-license records.
Documentation and support
Technical documentation, training/support and warranty obligations
Manuals, deployment runbook, handover exercise, support process and warranty terms.
Use the full SOW as the authority for exact wording; this table is a reading aid. The milestones page captures the concrete schedule terms.
Why local inference is the proposed default
The ownership and non-subscription language favors a system that can keep running without a required commercial speech API. The proposal is to deliver application code, approved model artifacts, configurations, and deployment machinery into the Library's environment.
That is an architectural interpretation, not a finding that every hosted service is forbidden or that open weights automatically satisfy the contract. AWS infrastructure still costs money. Third-party code, model weights, voice recordings, and commercial components each need their own rights assessment. Library ownership of delivered work does not erase third-party licenses.
Scope boundaries to settle
Agree which languages, book types, complex tables, equations, poetry, footnotes, and mixed-language passages are included in acceptance. Agree whether page navigation is required only when authoritative page references exist. Decide whether “various multimedia types” means an initial audio-plus-navigation product or additional full-text outputs. See open decisions.
A practical traceability record
For each accepted requirement, record its source section and edition, implementation component, test procedure, evidence artifact, and acceptance status. Keep this separate from model confidence scores. A model saying that a book looks correct is not proof of contractual acceptance.
Example proposal: the requirement for a valid navigation target maps to the compiler's reference checker, a deliberate broken-target test, and a playback test that reaches the expected chapter. This connects a rule to observable behavior.
The supplied materials do not identify one unambiguous standards baseline. The draft repeatedly references Z39.86-2002, while the supplied ANSI/NISO attachment is the 2005 edition reaffirmed in 2012. Record and resolve this conflict before building an exporter.
Editions in the source package
Source
Identified edition
How to use it
Draft SOW
ANSI/NISO Z39.86-2002; applicable NLS specifications, with current-version language
Starting statement of requested work; request an explicit precedence ruling.
Comparison reference; not an automatic replacement for J1.
The public 1203:2025 specification still references the 2002 standard and removes the older 250-file limit. The attached edition and the public edition therefore need separate profiles until NLS confirms which applies. Public construction specification.
Why the edition changes code
An edition can change element names, allowed structures, metadata, counts, or validation behavior. For example, the 2005 NCX has pageList and pageTarget; the 2002 standard uses navList and navTarget for page navigation. A generic DAISY exporter can produce valid output for its own target and still miss the required NLS target. Compare the 2002 standard with the attached 2005 standard.
The SOW's reference to multiple multimedia types also needs clarification against J1's audioNCX metadata requirement (§5.3). The proposal is to deliver the approved audio-and-navigation profile first while preserving internal text structure for possible future outputs.
Proposed versioned profile
Treat an output profile as configuration with an immutable version, not scattered constants throughout the code. It should identify:
Governing standards and exact schema/DTD/entity files, with hashes.
Required metadata, navigation classes, announcements, and supported multimedia types.
Master-audio and compressed-audio parameters, timing rules, and container constraints.
File names, file-size/count limits, checksum format, protection rules, and ZIP behavior.
Approved validator versions and target playback systems.
Documented NLS clarifications or authorized deviations, including their scope.
Every build records its profile ID. A change in profile creates a new build lineage so that later reviewers can explain why two otherwise similar books differ.
A profile selection process
Send NLS a short matrix of conflicting requirements, with source sections.
Ask for written confirmation of editions, precedence, output types, and supplied tools.
Turn the approved rules into a machine-readable profile and human-readable checklist.
Compile a small reference book and test it with the approved tools and players.
Freeze that profile for the pilot; handle later amendments as explicit version changes.
These are proposed implementation steps. The wiki does not establish legal or contractual precedence between conflicting documents.
Current references are separate from the baseline
The public mastering, protection, and delivery pages should be checked alongside the approved profile. Preserve dated copies or hashes when implementation begins; a live link may change after this wiki is written.
The proposed system has two front ends and one production core. Electronic text needs parsing and narration. Existing audiobooks need structure recovery and supplemental speech. Both converge on the same book representation, audio processing, compiler, validation, and delivery stages.
How the parts connect
Both input routes preserve their source evidence and feed the same book model, compiler and acceptance process.
Proposed architecture. The input route changes preparation and speech generation; both routes use the same controlled construction and acceptance process. Failed checks return to the affected stage for repair.
Who does what
Work
Proposed approach
Reason
Recognize uncertain structure and ambiguous text
Deterministic parsing first, model-assisted proposals when useful
Source evidence stays available and uncertainty can be reviewed.
Generate narration and announcements
Approved, replaceable speech models
Voices and model families can be evaluated without rewriting the compiler.
Maintain IDs, source coverage, references and timing
Typed data structures and conventional code
These properties need exact, repeatable checks.
Construct XML, encode, protect and package
Versioned deterministic compiler plus approved media tools
Compatibility depends on exact rules and artifacts.
Accept a book
Agreed validators, playback tests and authorized review
Passing a model check is only one source of evidence.
The main modules
Ingestion stores the original files and parses format-specific structure. The book service manages immutable versions, source spans, metadata claims and approvals. Speech and analysis workers synthesize or inspect bounded segments. The compiler uses accepted content and a frozen timeline to emit the required files. Independent validators report defects. The review interface lets an operator resolve exceptions and trigger dependent rebuilds.
The delivery adapter publishes only accepted, complete outputs through the agreed NLS interface. It should not infer an interface from the fact that the environment uses AWS.
Separate three kinds of segmentation
A speech chunk is a unit that can be regenerated. A navigation element is a destination a reader can select. A delivered audio file is a packaging and playback unit. They do not need to have matching boundaries.
For illustration, one chapter could contain 80 synthesis chunks, 12 navigable sections, and one delivered chapter audio file. Those numbers are explanatory, not requirements. Keeping the concepts separate enables local repairs while respecting file limits and meaningful navigation.
Illustrative chapter · counts from the example below · not a timing scale
Three kinds of segmentation solve different problems. A correction can replace a small speech chunk while navigation and delivery remain organized at larger scales. The counts are illustrative, not NLS limits.
A book behaves like an incremental build
A pronunciation correction can affect several speech chunks. Regenerate those chunks, reassemble the affected masters, recalculate the dependent timeline, rebuild navigation/XML, encode and protect the changed content, then update checksums and packaging. Unaffected accepted synthesis can stay cached.
This is a dependency graph: each output records the exact inputs from which it was made. If an input changes, downstream evidence becomes stale. The build system should never keep a green validation result for an old version of a file.
Boundaries for AI assistance
Model outputs are proposals or generated media until verified. They must not silently delete source text, invent page numbers, substitute a narrator, write arbitrary final XML, or mark their own output accepted. Books and metadata are untrusted content, including instructions that appear inside them.
A useful model response is structured: a proposed chapter boundary, the source span supporting it, alternative interpretations, and confidence evidence. The compiler consumes an accepted decision, not a free-form conversation.
Small services before many services
The module boundaries can initially live in one application with separate worker processes. Splitting every module into a network service is not necessary to prove the approach. Establish stable data contracts, job IDs and reproducible artifacts first; deploy independently where isolation or scale justifies it.
The canonical book model is the system's shared record of what the book contains and how each part becomes audio and navigation. “Canonical” means that downstream stages refer to the same structured representation; it does not mean that an AI transcript replaces the original source.
Proposed invariant: every source passage is represented in the output or has an explicit, reviewable reason for exclusion. This is the strongest protection against quiet omission, duplication and paraphrasing.
Follow one passage through the system
Proposed data model · trace one passage through a build
Source passageOriginal text + stable location$1,205
→
Approved spoken text“one thousand two hundred five dollars”
First arrow: retain the normalization rule and approval. Second arrow: retain the model, voice and lexicon versions.
Source span + stable node IDPreserves hierarchy and reading order
→
Structure recordHeading, paragraph, note, table or page marker; parent + order
Structure + accepted audioResolves logical destinations into media
→
Two navigation referencesMain narration destination and separate spoken heading-label range
Approved text + observed audioASR is evidence of what a recognizer heard
→
Quality recordArtifact version, check, evidence, defect and resolution; ASR stays separate from approved text
Exact input versions + dependent checks↓
Auditable output buildProfile and tool versions, dependency records and package hashes identify exactly what was delivered.
One passage, connected records. This proposed model preserves the source while linking each transformation, media asset and approval to an exact version. The currency example illustrates normalization; it is not an NLS-prescribed phrase.
What to retain
Record
Essential information
Why it matters
Source artifact
Original bytes, hash, format, ingestion version
Reproduce and audit a conversion.
Source span
File and stable structural location; original text or audio range
Locate the evidence behind every transformation.
Structure node
Heading, paragraph, note, list, table, page marker; parent and order
Preserve reading order and navigation semantics.
Text transformation
Original text, approved spoken form, rule, rationale, reviewer
Distinguish faithful normalization from accidental rewriting.
This schema is a proposal. Its final field names and storage technology are implementation choices.
Preserve three forms of text
Store original written text, approved spoken text, and observed ASR text separately. Each answers a different question: what did the source say, what should be spoken, and what did a recognizer hear?
Example:
Original text: $1,205
Spoken text: one thousand two hundred five dollars
Transformation: currency/USD
Source location: chapter 4, paragraph 12, characters 18–24
ASR evidence: populated only after examining the generated audio
A normalization rule should never convert “fifty” and “fifteen” into the same comparison token. Preserve high-consequence details such as negation, units, dates and proper names.
Use stable IDs and immutable versions
Give structural nodes stable identities across edits when they represent the same passage. Give source snapshots and generated assets immutable version IDs or content hashes. A node can point to a new approved audio asset without overwriting the old one.
A timestamp alone is an inadequate identity: earlier edits can shift every later timestamp. Store the logical destination and the resolved audio position, then derive the latter for each completed build.
Track two independent audio references
A navigation item points to a destination in the main narration and may also reference a spoken label in a separate headings file. Record both asset identities and timing ranges explicitly. Confusing the two timelines can produce a correct-sounding label that jumps to the wrong place.
Explore which work changes
Proposed dependency model · explore a change
A lexicon correction regenerates affected speech. Downstream timing and packages follow the changed audio; unrelated accepted synthesis can be reused.
Rebuild or rerunReuseDepends on scope
1
Original source + structureReuseRetain the source snapshot and stable passage identities.
↓ source spans + approved spoken forms
2
Main narration chunksRebuild affectedRegenerate passages using the corrected lexicon entry; reuse unrelated accepted chunks.
Spoken heading labelsDepends on scopeRegenerate only labels that use the corrected entry.
↓ accepted media versions from both branches
3
Assembled audio + final timingRebuild affectedReassemble affected masters and calculate final sample positions.
↓ resolved destinations + clip ranges
4
Navigation + XMLRebuild dependentRebuild dependent SMIL/NCX references and durations from the new timeline.
↓ compiled content + output profile
5
Encode, protect + packageRebuild changedProcess changed media and update dependent protection, checksums and packages.
↓ exact new artifact hashes
6
Validation evidenceRerun dependentValidate the changed dependencies and completed build; prior approvals remain tied to prior hashes.
Scope matters: a surname can have different pronunciations in different contexts. Record which occurrences the accepted correction covers.
A local edit can reach the final package. This simplified dependency order illustrates the proposed rebuild strategy. Narration and heading labels have separate assets and timing references. The approved profile determines whether a heading clip is copied or synthesized. The implementation must follow the actual dependency graph.
Worked repair example
An operator corrects a surname in the book lexicon.
Dependency lookup finds the passages whose spoken forms use that entry.
Those passages receive new speech recipes and audio versions.
The affected assembled masters and their final sample positions change.
Dependent SMIL/NCX timing, durations, checksums, protection and packages are rebuilt.
Validation runs against the new artifact hashes; the prior build remains auditable.
The system should record who accepted the correction, why it was made, and whether all occurrences should change. A name may intentionally have different pronunciations in different contexts.
Define cache identity carefully
Include source/spoken text, model and voice artifacts, normalization and lexicon versions, generation parameters, and relevant processing versions. Store the accepted audio itself. A random seed cannot guarantee byte-identical future synthesis across different hardware or software versions.
Ingestion should retain source evidence before attempting to improve it. The workflow differs by input format, but every importer should produce source-linked content and proposed navigation for the common book model.
Visual styling that disguises semantic structure; broken reading order.
Word
Heading styles, outline levels, lists, tables, footnotes and endnotes
Direct formatting that looks like a heading but has no heading style.
RTF or legacy Word
Convert in an isolated process while keeping the original and a conversion record
Lost notes, tables, symbols, or outline information.
Plain text
Delimiters and deterministic heuristics, followed by model-assisted classification
Missing hierarchy and ambiguous headings; uncertain reading order.
The format list comes from the SOW §§3.1–3.2; the importer design is proposed. Do not flatten a structured source into plain text before extracting its relationships.
Main content determines navigation
J3 §3.1 directs navigation work to the main content rather than simply copying the table of contents. A table of contents may omit subheadings or describe them differently. Inspect the actual text and preserve evidence for every proposed division.
J3 also separates hierarchical navigation from nonhierarchical structures such as notes and segments. A note is not necessarily another chapter. The correct distinction affects playback behavior and must survive ingestion into the compiler.
A model can propose that a passage is a section heading, but the proposal should include its source location and supporting evidence. Printed page numbers should come from an authoritative page map or source, not model invention.
Commercial recordings
Use evidence in this proposed order: publisher metadata and chapter boundaries, matching electronic text when available, an ASR transcript, semantic classification, and precise audio alignment. Record conflicts rather than silently choosing whichever source arrived last.
For example, a model can distinguish a spoken chapter announcement from dialogue mentioning “chapter three.” The final destination still needs alignment and a waveform-level boundary check.
Keep the supplied narration through approved mastering and conversion. The ASR transcript is an analysis aid that can contain errors. Synthesize only the additional speech the approved output requires. Availability of a narrator's recording is not permission to clone the narrator's voice.
A recording cannot reliably reveal an unspoken print heading or an unannounced print page number. Mark these as unavailable, and agree how that input class will be accepted.
Written text to spoken text
Normalization expands abbreviations, dates, numbers and symbols while retaining original source spans. The sentence “Dr. Smith paid $1,205 on 03/04/2026” needs context: “Dr.” may require disambiguation, the currency needs expansion, and the date convention is ambiguous.
Use deterministic rules for well-defined cases and a reviewable proposal for contextual cases. A book-level pronunciation lexicon keeps repeated terms consistent. NVIDIA NeMo text processing is a candidate foundation for text normalization, not a guarantee that the whole book is correctly interpreted.
Model adapters and source safety
Each speech adapter should declare which controls it supports: SSML, phonemes, pronunciation dictionaries, language selection, reference audio, and style instructions. Unsupported controls must fail clearly or use an agreed conversion; they should not be silently ignored.
Treat model-specific control strings in source text as content. Escape or otherwise neutralize them before a model can interpret them as a voice or performance instruction. Retain the literal source and record the transformation.
What an importer should report
Proposed outputs include an ordered content tree, source coverage map, metadata conflicts, lost-format warnings, navigation proposals and unresolved passages. A successful file parse alone is insufficient: compare recovered content and reading order with the source before moving into speech production.
Use a replaceable text-to-speech (TTS) interface and evaluate several local models on the same material. The names below are candidates from the original analysis. None has been demonstrated here to satisfy NLS acceptance requirements.
Model-card licensing requires attention, including separate commercial licensing
Conditional evaluation candidate after terms are resolved.
License labels do not establish rights to every associated voice, dataset, dependency or deployment use. Pin the exact artifacts and record applicable terms in the handover inventory. This wiki's model descriptions are not a legal license clearance.
A book level narration plan
Before synthesis, select an approved voice, pronunciation lexicon, language handling, and restrained pacing/style settings. Preserve these across chunks. Faithful reading is the default: selectable voices do not imply automatic character acting, added laughs, paraphrasing or sound effects.
Test paragraph or sentence-group chunks. The original analysis suggests experimenting with 20–60 seconds of speech per chunk; this is a tuning range, not a requirement or measured optimum. Preserve sentence boundaries, and compare voice continuity, pronunciation, joins, and correction cost at different lengths.
Bounded generation and repair
Synthesize a chunk from approved spoken text and a recorded recipe.
Transcribe with an independent recognizer and compare meaningful content.
Check alignment, clipping, silence, level and joins.
Repair a specific discrepancy through a lexicon, text, boundary or generation change.
Escalate after an agreed attempt limit rather than retry indefinitely.
Do not silently switch narrators for a difficult sentence. An engine fallback must preserve the approved voice or trigger a deliberate review and potentially a larger regeneration.
ASR and alignment candidates
Qwen3 ASR provides transcription and a separate forced aligner. WhisperX combines transcription with alignment and documents limitations. Long books need bounded jobs with reconciled overlap, rather than a single whole-book alignment call.
ASR asks what a recognizer hears. Forced alignment asks where expected text fits the audio. These are different checks, and neither establishes perfect narration alone. See the validation page.
A useful evaluation corpus
Include ordinary prose and long chapters alongside names, dates, quantities, negation, abbreviations, dialogue, notes, lists, tables, mixed-language passages, and repeated terms. Add passages from each supported input class. Keep tuning material separate from a held-out acceptance set.
Measure complete books as well as excerpts. Report unresolved defects per accepted audio hour, review minutes, repair attempts, throughput, and consistency across joins. Do not infer whole-book quality from a 30-second demo or time to first audio.
Voice governance
Keep approved voice assets, consent or usage rights, version, language scope and allowed transformations with the production recipe. A change to the voice reference can affect every generated passage and therefore should invalidate the appropriate cache and quality evidence.
The compiler converts an accepted book model and finalized audio into the files required by an approved NLS output profile. It should serialize typed data and enforce explicit rules. Generating final XML directly with an LLM would make exact compliance harder to control.
What the main files do
File or format
Role
OPF
Describes metadata, the package contents and reading/playback organization.
NCX
Defines navigable structures, labels and destinations.
SMIL
Connects ordered playback structures to particular media clips.
WAV masters
Accepted uncompressed source audio used as the timing reference.
AMR-WB+ in 3GP
Compressed audio with NLS-specific codec and container constraints.
Resource files
Supply approved alternative navigation-class labels when needed.
dtb.md5
XML checksum document for the DTB files, excluding itself.
DTD and entity files
Local dependencies needed to interpret and validate the XML.
These descriptions summarize J1 §§3.1–3.10. Exact applicability depends on the selected profile.
Freeze audio before final timing
Proposal: track positions as integer samples, complete approved mastering and assembly, then freeze the accepted WAV timeline. Derive clip timestamps from that final timeline, not from rough ASR estimates or an earlier synthesis chunk.
In attached J1 §3.2.5, timestamps remain relative to the accepted WAV; codec offset compensation is assigned to the player. Do not silently shift XML to encoded-audio timing or compensate twice. J1 pp10–11.
Timing windows depend on clip type
Hatched band = permitted location of the clip boundary
Main SMIL narration + heading labels
Both clip boundaries must fall in silence.
Before speech starts
−200 ms0
Start: 80–120 ms before
Speech
duration schematic
After speech ends
0+500 ms
End: 150–300 ms after
Alternative-class resource audio
Use the resource-specific window; its allowed range differs.
Before speech starts
−200 ms0
Start: 100–200 ms before
Speech
duration schematic
After speech ends
0+500 ms
End: 150–500 ms after
Read from the speech edges outward. Each zero marks the adjacent speech boundary. Lead-in and tail use separate scales; the speech block has no time scale. Select one endpoint within each permitted band. These are boundary windows, not required amounts of inserted silence. If a valid boundary is unavailable, flag the case or approve an audio edit and rebuild timing.
The first two rows also require the boundaries to lie in silence. Both clip endpoints must be present. A shared global “navigation padding” constant would miss the resource exception.
Use alignment plus local waveform analysis to identify the actual speech boundary. If the required margin is unavailable, flag the case or make an approved audio edit and rebuild dependent timing. Choosing the middle of a permitted window is a proposed default, not an NLS rule.
Main destinations and spoken labels
A navigation destination resolves into primary audio through SMIL. Its spoken label may refer to a different range in the headings audio. Keep both references explicit. Labels must represent the heading, not the whole section. J1 §§3.4.3.4–3.4.3.5, pp13–14; J3 §3.1.2, p6.
J3 requires at least two level-one navPoints: the first marks the opening title/author with class title/author, and the last marks the closing announcement with class close. J3 §3.1.1, p5.
Notes, note references, sidebars, pages, line numbers and segments need their own semantics. J3 requires segment navigation for entries in specified lists, and J1 says not to create navLists for segments. Treating every navigation object as a chapter loses these distinctions. J3 §§3.2.1–3.2.4, pp6–7; J1 §3.4.4, p14.
Exact rules from the attached construction edition
Rule in J1
Compiler consequence
No more than 250 files; §3.1.3
Count all required files under that edition. The newer public edition differs.
Split SMIL above 100 KiB; at most 50 SMIL files; §3.3.12
Measure actual UTF-8 serialized bytes; 100 KiB is 102,400 bytes. Follow the stated escalation rule if limits are exceeded.
At most 5,000 navPoints; contact monitor when exceeding 1,000; §3.4.3.6
Distinguish a hard limit from a consultation trigger.
Lowercase filenames and prescribed production IDs/suffixes; §3.1
Generate names from the profile and NLS-supplied identity.
XML diskcheck checksum structure; §3.9
Ordinary md5sum text is insufficient. Include each file except the checksum itself; do not list the checksum in the OPF manifest.
Include referenced DTDs and entities; UTF-8 XML; §3.10
Package dependencies and validate against an approved local catalog.
Duration follows SMIL playback, including specified custom-test content; §5.3(v)
Derive duration from the compiled playback structure, within the stated ±1 second accuracy.
See J1 pp9–20. Do not apply the attached 250-file limit to a newer profile by accident. Permit only approved local XML dependencies; do not enable arbitrary network/entity resolution for book content.
Encoding is an early interoperability risk
AMR-WB+ is not ordinary AMR-WB. Attached J1 specifies mono constant bitrate, frame type 23 and ISF index 8, plus particular 3GP metadata and field constraints. A file extension alone proves none of those properties. J1 §3.2.3, p10.
Hindenburg's support article describes an NLS-supplied encoder; APH's revision history records NLS protection/validation integration. These are historical precedents. Obtain current components, terms, supported operating systems, automation interfaces and cloud-deployment approval before relying on them.
Protection is part of the book
PDTB protection involves protected content and associated authorization/key material. S3 storage encryption protects stored objects but does not create a protected talking book that an NLS player can authorize. See the DAISY protection specification and request the governing NLS 1205 profile and tooling.
Packaging and final submission
The public 1206:2025 delivery specification describes an outer ZIP with protected/unprotected DTB archives and WAV masters. It specifies uncompressed entries and local-header CRC/size values with ZIP bit 3 clear. A seekable staging file is a proposed way to control those headers, followed by archive inspection and extracted-file checks. Delivery specification §§3–3.2.
Its §4 names the NLS Transfer Portal. Internal S3 staging is a separate architectural choice; confirm the project-specific final interface. Keep QA logs outside the delivery package unless that interface expressly permits them.
Reuse existing DAISY components selectively
DAISY Pipeline's DTBook-to-DAISY3 script accepts a 2005-3 DTBook input and offers speech-related conversion options. It is a useful candidate for selected transformations. That documentation does not establish an NLS 2002, AMR-WB+, protected export path. Keep the NLS profile and acceptance tests explicit even when reusing its code.
Automation succeeds when defects are detectable and repairs are precise. A generated file is only an intermediate result until content, audio, navigation, format and playback checks agree with the accepted requirements.
The following layered QA design is proposed. The draft requires quality and compatibility testing but does not provide numerical narration-error or review-effort thresholds. Those must be agreed before the pilot is judged. SOW §§3.5, 3.11, Phase II.
From a candidate to an accepted release
Generate a candidateBook model, narration, navigation and package
↓ candidate + artifact hashes
Check six layersSource · speech · acoustics · navigation · format · playbackModel checks are evidence, not acceptance.
↙ issues found
Bounded repair loop
Triage the findingIdentify the cause and the smallest valid repair.
↓ repair is clear; budget remains
Repair affected artifactsRegenerate, correct or make an approved edit.
↶ Return to the six checksRecheck every result invalidated by the repair, including dependent timing.
Ambiguous issue or retry budget spent?→ Escalate to an authorized reviewer. Record the decision and evidence.
↘ no unresolved blockers
Separate release decision
Final acceptance gateApply the agreed content, audio, navigation and delivery criteria.Review evidence independently and verify the agreed player/app matrix.Document any authorized exceptions.
↓ accepted against agreed criteria
Accepted release
← Rejected: return to triage; preserve the failed build and findings.
Check results belong to a specific artifact hash and output profile. A repair creates a new candidate; earlier passing results may no longer apply.
Proposed control flow. The retry budget, acceptance criteria and review authority must be agreed. The draft calls for quality and compatibility testing; the numerical acceptance thresholds still need agreement. Draft SOW §§3.5, 3.11 and Phase II.
Six layers of correctness
Layer
Example defect
Proposed evidence
Source fidelity
A paragraph is omitted or reordered
Source-span coverage and ordering checks against the imported book.
Speech fidelity
A number changes or a word is repeated
Independent transcription, protected-token comparison and targeted listening.
Acoustic quality
Clipping, an abrupt join, long silence or voice drift
Signal checks plus excerpt and sustained-listening review.
Navigation
A label announces one chapter and lands in another
Reference resolution, alignment and interaction tests.
Format
Broken XML references, wrong codec, bad checksum or ZIP headers
Schema, semantic, media and package inspectors.
Playback
Protection, skip, resume or speed behavior fails on a target player
Recorded test matrix using actual agreed players/apps.
ASR evidence is not ground truth
An ASR mismatch can arise from a TTS mistake, a recognizer mistake, or a legitimate difference between written and spoken forms. Forced alignment fits expected text to an audio signal; it cannot independently prove every expected word was spoken.
Combine independent transcription, alignment, acoustic evidence and targeted checks for names, quantities, units, dates and negation. For high-consequence discrepancies, use another recognizer or a human listener. Agreement between models increases evidence but does not guarantee correctness. WhisperX's documented limitations illustrate missing timestamps and other alignment problems.
A review screen should show enough evidence
For each issue, present the source passage, approved spoken text, recognizer output, short playable excerpt, timeline/waveform, reason for the flag, affected build and repair history. Offer concrete actions: correct pronunciation, edit a proposed boundary, regenerate a passage, or accept a documented exception.
An exception should retain rationale and authorized reviewer identity. Acceptance of one passage should not automatically suppress the same error class throughout a book.
Defect states and bounded retries
A proposed workflow is detected → triaged → repair attempted → rechecked → accepted or escalated. Set an agreed retry budget per issue. Distinguish processing failures, which may be safely retried, from content ambiguities, which often need a decision.
Retest everything invalidated by a repair. Changing a chunk's duration can invalidate later clip times even when the later audio remains unchanged. Tie every check result to the exact artifact hashes and output profile.
What to measure
Metric
Definition to agree
Why it helps
Unresolved content defects
Confirmed omissions, substitutions or repetitions per accepted audio hour
Tracks fidelity more directly than general “naturalness.”
Navigation accuracy
Correct labels and destinations divided by evaluated items, by type
Prevents chapter success from masking note/page failures.
Review effort
Active human review minutes per accepted audio hour
Measures the practical cost of automation.
Repair burden
Attempts and regenerated audio per accepted hour
Exposes unstable models and costly chunk sizes.
Throughput
Accepted output hours per elapsed production hour and per compute hour
Separates machine speed from accepted-book delivery.
Operational reliability
Failed/resumed jobs, incomplete packages and recovery time
Tests production behavior rather than isolated demos.
These are proposed definitions, not measured results or contractual targets. Report the test corpus, input class, language, hardware, model version and accepted-output denominator with every result. Avoid hiding rejected books from quality statistics; report rejections separately.
Playback tests beyond the opening chapter
Exercise all navigation levels, notes/sidebars and skipping behavior, segment stepping, resume, transitions between audio files, and faster playback. Check both protected and unprotected artifacts where relevant. Record the player/app version and profile.
The draft expressly calls for testing with the Library playback environment. The exact device matrix and pass criteria remain open decisions. Accessibility of the operator interface should also be tested with keyboard and assistive technology as a proposed implementation practice.
Keep acceptance independent
Use known reference books, deliberate malformed examples, and a held-out evaluation set. An AI coding assistant can help generate tests and analyze regressions, but the acceptance criteria and reference outputs should be reviewed independently from the implementation that produces them.
The draft requires integration with the Library's AWS systems and standard import/export without proprietary lock-in. The components below are a proposed implementation, not a mandated stack. Confirm available services, network boundaries, security rules and delivery interfaces with the Library. SOW §3.7.
Proposed components in the Library’s approved AWS environment
Application + control planeBook versions, jobs, approvals and output profiles
↔state
PostgreSQLStores state and decisions for the control plane
↓ schedules jobs and records results
Workflow orchestrationStep Functions + AWS Batch, or an approved equivalent
↓ dispatches bounded, resumable work
Model workersTTS · transcription · alignment Pinned models and voices; GPU or suitable CPU
Compilation workersParsing · XML · media tools · validation · packaging Isolated CPU jobs
↕ workers read inputs and write versioned artifacts
S3 artifact storage
Originals→Intermediates→Staging→Accepted outputs
Delivery adapter Submits only the accepted package; records acknowledgement
↓ crosses the agreed delivery interface
NLS receiving workflowPublic 1206 guidance names the Transfer Portal. The project-specific integration still needs confirmation.
Proposed deployment, not a prescribed AWS stack. The control plane manages decisions, workers process content, S3 stores artifacts, and a separate delivery adapter submits accepted books. S3 storage is not itself the confirmed submission interface. The diagram shows logical responsibilities, not a network-security design.
Proposed deployment components
Component
Candidate implementation
Responsibility
Control plane
Typed application service
Book versions, jobs, profiles, approvals and audit events.
Metadata and job state
PostgreSQL
Structured records, transactions and searchable status.
Artifact storage
S3
Separate originals, intermediates, staging and accepted outputs.
Speech and analysis workers
Containerized Python workers with pinned models
GPU inference or appropriate CPU inference.
Compilation workers
Isolated CPU containers
Parsing, XML, media inspection and packaging.
Workflow orchestration
Step Functions with AWS Batch, or approved equivalent
Dependencies, bounded retries, checkpoints and review gates.
Operator application
Accessible web interface
Exception review, progress and authorized release.
Delivery adapter
Agreed NLS integration
Submit validated packages and record acknowledgement.
Use immutable inputs, explicit output versions and idempotent job keys. A retried job should not duplicate a submission or overwrite an accepted book. Record checkpoints after durable artifacts are written and verified. Use separate states for failed processing, unresolved content, ready for review, and ready for delivery.
Keep model workers warm over a batch where appropriate. Loading weights for every paragraph wastes time, but the right worker lifetime depends on measured throughput and utilization.
Capacity without invented benchmarks
A useful planning equation is:
GPU-hours of work = source audio-hours × processing factor × retry multiplier
Wall-clock hours ≈ GPU-hours of work ÷ (GPU count × effective utilization)
Here the processing factor must be measured in GPU-hours per source audio-hour for the chosen workload. If it is a combined factor, avoid counting transcription or repair twice.
For illustration only: 30,000 hours × 0.5 GPU-hours per hour = 15,000 GPU-hours before retries. At a 1.2 retry multiplier, this becomes 18,000. Ten GPUs at 75% effective utilization would imply about 2,400 elapsed hours of compute processing. These are assumptions and arithmetic, not observed performance or a calendar delivery forecast. Human review, CPU work, queues and delivery can remain bottlenecks.
Benchmark the two input paths separately. Commercial recordings may need full-book transcription/alignment but only supplemental TTS; electronic text needs full narration. Measure accepted output, including failed attempts, rather than quoting isolated generation speed.
Security and source handling
The draft references federal controls in Section H, which is absent from the supplied attachments. It also calls for source/output protection, retention/destruction procedures and security documentation. Obtain that missing material before fixing the controls baseline. SOW §§3.6, 3.13.1, 5.1.
Proposed defaults include isolated inference in the approved environment, restricted unnecessary egress, pinned dependencies, least-privilege roles, separate content-protection keys, and limited source text in logs. Book content must be treated as data, including text that looks like instructions. SOW §5.1 limits use of Government source materials to contract performance and requires secure disposal at contract end. Implement those obligations alongside the agreed retention and deletion procedures.
Storage encryption, transfer encryption and PDTB protection solve different problems. Keep their responsibilities and key handling explicit.
Handover and maintenance
Deliver source code, deployment instructions, configuration, dependency/model/voice inventory, operator and administrator guides, recovery procedures, validation tools and a known reference build. Confirm which third-party components can be redistributed and which must be supplied by NLS.
The draft specifies support response within 24 hours and a 12-month defect-correction warranty after acceptance; it permits ongoing support/maintenance separately from prohibited recurring access requirements. See the milestone page and SOW §§3.8–3.9, 3.12.
Release evidence
For each released software/profile version, preserve migration notes, a reproducible deployment recipe, dependency hashes, regression results and the reference-book outputs. That makes later maintenance possible without relying on the original contractor's environment.
Use technical proof gates to reduce uncertainty while meeting the draft's separate contractual phases. The five gates below come from the engineering proposal; they are not the four phases named in the SOW.
Proposed proof gates
Gate
Work to demonstrate
Pass evidence
1. Format and protection
Build a tiny complete book with the approved profile, encoder, protection and packaging
Approved validation results and working navigation/playback on target systems.
2. Long-form speech
Compare candidate engines on chapters and complete representative books
Agreed fidelity and quality results, with measured correction effort.
3. Both ingestion paths
Process structured text, less-structured text, commercial audio with and without matching text
Source coverage and independently checked navigation, including explicit unavailable information.
4. Repair and rebuilds
Change pronunciation, replace a paragraph, edit a heading, interrupt processing and fail encoding
Only necessary work repeats; all dependent timing, hashes, metadata and packaging remain correct.
5. Production operation
Run batches, security checks, recovery and real integration
Repeatable accepted outputs, measured throughput and operational handover.
Prove codec/protection compatibility early with simple known content so that format errors can be separated from speech errors. Speech and importer experiments can proceed in parallel after their required inputs are available.
Schedule in the supplied draft
POP means period of performance. No actual start date or total duration is supplied here, so the dates below are relative triggers rather than calendar commitments. SOW §4.1 and named phase descriptions.
Deliverable or event
Draft timing
Key personnel resumes
At proposal submission.
Kickoff
Within 10 business days; the table uses POP start while §5.2 uses contract award. Clarify the trigger.
Five architecture diagrams and eight technical documents
Within 30 calendar days after POP start.
Status and quality-control reports
Every 30 calendar days during POP.
Phase I prototype, sample DTBs, workflow demonstration and test report
Within six months of POP start.
Phase II integration and operational testing
Starts after Phase I completion and Government acceptance; completed within 90 calendar days.
Phase III full deployment
Starts after Phase II completion and Government acceptance; completed within 90 calendar days.
Phase IV ongoing support
After Phase III completion and acceptance, through the remainder of POP.
Warranty
12 months after written Government acceptance of all deliverables.
The 24-hour support-response requirement appears in §3.8. The warranty includes free defect correction and patches under §3.9. Do not turn the phase durations into a promised total project length: Government acceptance timing and the total POP are not established here.
How the gates can fit the phases
Proposed mapping: use Phase I to resolve the profile and prove format/protection, model behavior and initial ingestion. Continue hardening ingestion, repair and operational behavior through Phase II integration tests. Phase III adds deployment, handover and acceptance evidence. Phase IV covers production support and controlled improvement.
This mapping does not reduce the deliverables required in any phase. Validate it against the final solicitation and approved project plan.
Early documents should expose real decisions
The first-month architecture package should show both source paths, data ownership, security boundaries, failure/retry handling, and the split between AI proposals and accepted compiler inputs. List unresolved standards and supplied-component dependencies rather than hiding them behind a generic cloud diagram.
Suggested readiness check before a pilot
Confirm the output profile; acquire reference books, target players and required tooling; agree supported content classes and quantitative acceptance criteria; record authorized voices and dependency rights; obtain Section H; and identify who can accept exceptions and release a book. These are suggested prerequisites tied to the decision register.
The items below remain unresolved by the supplied materials. Suggested decision partners describe the kind of authority needed; they are not assignments made by NLS. Do not mark an item resolved without a written answer, accepted test result or other attributable evidence.
Decisions that affect the architecture
Question
Why it matters
Evidence needed and suggested decision partner
Which editions and precedence govern the 2002/2005 conflict and current-version language?
Changes XML structures, limits and test rules.
Approved profile matrix from the contracting/technical authority.
Which multimedia types are required at acceptance?
audioNCX and synchronized full text imply different outputs.
Supported output-type list and reference books from NLS production/QA.
Which encoder, protector, validators and keys will NLS provide?
Determines the feasible encoding/protection route and deployment platform.
Current packages, versions, terms, headless interfaces, OS support and cloud approval.
What accompanies commercial recordings?
Metadata and matching text determine recoverable structure and pages.
Representative source sets and expected navigation from NLS production.
How do mastering rules apply to synthetic audio and generated labels?
Native model rate, initial-recording language and copied heading rules may conflict with a chosen workflow.
Approved interpretation and accepted synthetic example from NLS audio/QA.
What security controls and retention rules apply?
Determines network, access, key, logging and deletion design.
Missing Section H and approved architecture requirements from Library security.
What constitutes an accepted book?
Controls quality thresholds, human review and pricing/effort assumptions.
Content classes, language scope, player matrix, defect thresholds and throughput/turnaround targets.
What is the final submission interface for this project?
The public Transfer Portal and proposed internal S3 staging are different interfaces.
Integration contract, credentials/process ownership and acknowledgement semantics.
How are third-party model, voice and component rights handled?
Library ownership/non-subscription terms must coexist with component rights.
Approved inventory and applicable license/voice permissions from appropriate procurement/legal reviewers.
Specific mastering clarification
The live public mastering page is 1202:2026, revision 1.5, June 3, 2026. It specifies initial 44.1 kHz PCM and adds integrated loudness of −22 LUFS ±2; it also says heading clips are copied from primary audio. The public index still advertises a 2025 entry. Confirm the actual governing edition and how these rules apply to neural synthesis and separately generated navigation speech. Resampling a model's output alone does not settle the “initially recorded” interpretation. Current mastering page and specification index.
Additional source gaps
The SOW mentions 1201; J1 references 1208 for commercial-audio masters; only 1203 is among the four attachments. Request the complete applicable specification set. The phrase “and D standards” in SOW §3 is unexplained. The kickoff start trigger differs between §4.1 and §5.2. These are useful clarification items, not grounds to invent a preferred interpretation.
Proposed decision record
For each answer, store the exact question, governing source and section, answer date, responsible authority, evidence link, affected profile/version, affected tests and remaining conditions. Preserve superseded decisions as history.
Example: approving a new file-count rule updates the output profile and the corresponding validator. It should not quietly change an already submitted book or erase the rule used for its original build.
What can proceed while answers are pending
Build the source-preserving data model, importer experiments, independent quality checks and model evaluation harness. Prototype against a clearly labeled provisional profile. Avoid committing to a production codec/protection stack or promising numerical quality/throughput until the relevant components and acceptance rules are confirmed.
This glossary explains the terms as used in this proposed production system. The linked articles and standards contain the implementation detail.
Term
Meaning
NLS
The Library of Congress National Library Service for the Blind and Print Disabled.
DTB
Digital talking book: audio and supporting structure that allow accessible reading and navigation.
DAISY
Standards and technology for accessible publications, including structured talking books.
ANSI/NISO Z39.86
The talking-book standard cited in the procurement; the edition matters.
audioNCX
A multimedia type describing full audio with navigation structure; it does not by itself promise a synchronized full-text product.
TTS
Text to speech: generating spoken audio from text.
ASR
Automatic speech recognition: generating a transcript from audio.
Forced alignment
Estimating where a supplied transcript fits the audio; distinct from independently recognizing what was spoken.
Canonical book model
The shared structured record linking original sources, approved text, navigation, audio and evidence.
Provenance
Records of where an item came from and how it was transformed.
Source span
An identifiable passage or audio range in an original source.
Spoken form
The approved text representation intended for narration, such as expanding a currency amount.
Lexicon
A book-level or shared set of pronunciation decisions.
Inference
Running a trained model to generate or analyze content.
Model weights
Numerical parameters of a trained model; they are distinct from application code and voice assets.
Deterministic compiler
Conventional software that produces prescribed files from explicit accepted inputs and a versioned rule set.
Output profile
The exact standards, media, metadata and packaging rules chosen for a build.
OPF
The package document describing metadata, included resources and organization.
NCX
Navigation Control file for XML; defines navigation structures and labels.
SMIL
Synchronized Multimedia Integration Language; links playback structures to media segments.
DTD
Document Type Definition; a schema mechanism used by these XML formats.
Entity
An XML declaration/reference whose allowed definitions must be controlled and packaged as required.
navPoint
A hierarchical navigation point, such as a chapter.
navTarget
A target in an auxiliary navigation list; its allowed use depends on the standard edition.
Segment
A useful navigable unit that need not be a chapter or heading.
PCM / WAV
PCM represents uncompressed audio samples; WAV is a common container for those samples.
Sample position
An integer location in audio at a known sample rate, useful for precise internal timing.
AMR-WB+
The extended wideband audio codec required by the cited NLS profile; different from AMR-WB.
3GP
The container used for the compressed audio, with NLS-specific constraints.
LUFS
A unit for measured perceived loudness; the applicable mastering edition defines the target.
PDTB
Protected Digital Talking Book; content protection and authorization for the book/player ecosystem.
MD5
A checksum algorithm used in the specified DTB checksum format. A checksum is not an authorization or authenticity system.
CRC
A ZIP integrity check value; the delivery specification constrains how it appears in the headers.
Idempotent job
A job whose retry does not create unwanted duplicate effects.
Incremental build
Regenerate changed outputs and their dependencies while reusing unaffected accepted artifacts.
S3
AWS object storage; proposed for internal source and artifact storage here.
GPU
A processor commonly used to accelerate model inference.
Real-time factor
Processing time divided by audio duration under defined hardware and workload conditions.
POP
Period of performance: the contractual time window for performing the work.
SOW
Statement of work: the document defining requested work and deliverables.
RFI
Request for information: a market-research request, distinct from an award or accepted implementation.
For precise format definitions, use J2 or the 2002 standard according to the approved profile. General operational definitions in this table are explanatory additions.
Sources were checked for this local edition on October 1, 2026. Attachment citations link directly to the original SAM.gov downloads. Preserved local copies remain available as backups; live links may change. Nearby article citations identify the source of specific claims. Public references do not automatically supersede attachments or resolve procurement ambiguities.
Original opportunity and attachment context. Check it for later amendments.
The attachment links above are public SAM.gov download endpoints, without session tokens or API credentials. If a direct download becomes unavailable, return to the opportunity attachment list or use a local backup.
The original analysis and SOW text extract are local supplementary files, not attachments hosted on SAM.gov. The SOW text extract is a convenience copy; layout and table relationships may be flattened, so use the Word original to resolve ambiguities.
PDF links with #page= use one-based file page numbers, not necessarily the printed number on the page. Browser support for opening at a specific PDF page varies, and SAM.gov may download a PDF rather than display it at the requested page.
Catalogued November 2007 reference; direct retrieval failed during this review. Its clauses and governing revision were not independently verified here.
AWS Batch GPU jobs and Step Functions with Batch document orchestration options. The Library's approved environment and actual workloads determine whether those options fit.
What changed from the original analysis
The wiki reorganizes the analysis into linked topics and adds explanation rather than treating every recommendation as a requirement. Source review added resource-audio timing exceptions, WAV-relative timing semantics, exact attached-edition limits, draft milestones, support/warranty terms, a missing-specifications list, the live mastering edition, and the distinction between S3 staging and the public Transfer Portal.
The original analysis remains available unchanged for comparison. No benchmark, cost quote, accepted architecture, contract interpretation or playback demonstration has been added as a claimed result.