This wiki explains a proposed system for the Library of Congress National Library Service for the Blind and Print Disabled (NLS). It would turn electronic books and existing commercial audiobooks into accessible digital talking books (DTBs) with reliable navigation, required audio formats, protection, and delivery packaging.
The central recommendation is an AI-assisted production pipeline with a deterministic compiler at its center. AI proposes structure, prepares speech, and helps detect defects. Conventional software enforces the approved rules and creates the final files. The original analysis considers this feasible; no model benchmark, compliant prototype, or successful NLS-player demonstration is established by the materials here.
Source requirement means a statement found in the supplied draft SOW or a named specification. Edition and applicability still matter. Proposal means an engineering choice from the analysis or this wiki. Open decision means an answer is needed from NLS or from testing before the design can be fixed.
Public specifications are useful comparison material. A newer public edition does not by itself amend this opportunity. The standards page explains the 2002/2005 mismatch and the difference between the attached and current public construction specifications.
What to prove first
Produce a tiny, complete, protected book that passes the approved validators and plays correctly on the intended NLS systems. In parallel, test sustained narration on representative complete books and measure correction effort. These demonstrations address format interoperability and speech quality separately.
About this local edition
Prepared on October 1, 2026, from the supplied analysis, the four opportunity attachments, and linked primary sources. Added explanations include a glossary, explicit acceptance evidence, dependency tracking, milestones, and decision prompts. Nothing in this wiki records answers from NLS or actual performance results.
Wiki content works offline. Attachment citations open the original files on SAM.gov and require an internet connection. The source library also offers local backup copies. Editable articles are stored alongside it as Markdown. Keep the whole folder together when moving it. See the source library for provenance and the glossary for unfamiliar terms.
Proposed baseline: build a working, inspectable system that converts electronic text and supplied commercial audio into sample NLS digital talking books, with an operator able to correct a defect and rebuild the affected output. Start on an Apple Silicon Mac; prove the same workflow on one GPU in an approved AWS development environment. The proposed narration engine is Qwen3-TTS 1.7B CustomVoice, with Kokoro-82M as a compact comparison baseline; model selection remains conditional on listening tests and rights review.
Status: plan only. The wiki and its hosting are implemented; the production prototype, model benchmarks, purchased hardware, and NLS acceptance are not. Scope, staffing, durations, memory budgets and pass thresholds below are engineering proposals. Contract statements refer to the supplied draft SOW, not an awarded or complete contract.
The prototype should answer four questions: can the chosen voice read complete books faithfully; can the compiler produce a protected book that works on the intended players; can operators resolve exceptions without redoing a whole book; and can the Library operate the delivered system without a required paid speech API or recurring software-access license?
Phase I in the draft calls for an initial working TTS prototype, DTB samples, voice samples, an end-to-end workflow demonstration and a test report. The broader SOW covers both input routes. This proposed prototype deliberately exercises both so commercial-audio risks appear before Phase II. It does not imply that Phase I alone completes all integration, production or support obligations. Draft SOW §§3.1–3.2, 3.12–3.13; deliverable table §4.1.
Scope and boundaries
Area
Included in this prototype
Boundary or follow-on work
Text input
EPUB, DOCX, RTF and UTF-8 plain text; preserve reading order, headings, source spans and metadata
Confirm whether “Word Documents” includes legacy .doc; quarantine unsupported, malformed, encrypted or DRM-protected inputs rather than silently drop content. Scanned-page OCR is a separate extension.
Commercial audio
WAV, MP3, OPUS and FLAC fixture coverage; preserve narration, propose navigation from available metadata/text/ASR, synthesize approved announcements
Do not recreate the narrator or invent unavailable page references. Audio-only boundaries need review.
Languages and content
English prose first, two approved selectable voices; ordinary headings, lists, notes, numbers and abbreviations
English is a proposed pilot limit, not an SOW exemption. Mixed-language text, equations, poetry, complex tables and illustrations must be flagged and given an agreed reading treatment. Wider language support is a follow-on gate.
Output
One agreed audio-plus-navigation profile, WAV masters, required XML/metadata, AMR-WB+, PDTB protection and delivery package
MP3 is optional in the draft. Other multimedia types remain a scoped decision. A WAV/MP3 demo or unprotected package cannot substitute for required protected-output proof.
One production team and simple roles; no multi-tenant service, billing or public upload portal.
Automation
Three-book batch, durable checkpoints, bounded retries, cancellation, resumable work and dependency-based rebuilds
No promise to sustain the approximately 3,000-book annual background volume until representative benchmarks exist.
Deployment
Native Mac model worker, shared API/data contracts, Linux NVIDIA worker in approved AWS test environment
Real Library storage/distribution integration belongs to the agreed interface plan and Phase II acceptance. Prototype adapters may be simulated, visibly labeled.
Handover
Source, pinned models and voices where redistribution is permitted, deployment recipe, test evidence and operator/admin instructions
Production HA/DR, large-scale load qualification, full training and support remain later-phase work; backup/restore and access controls still receive prototype tests.
Each gate earns the next claim. Import and model experiments may run in parallel; missing protection tooling or player access leaves the corresponding output gate blocked.
Planning allowance: 16 engineering weeks, assuming roughly four full-time engineering contributors plus fractional project management/specialists, available source rights, an NLS technical contact and timely access to required tooling and players. This is an internal estimate, not a delivery promise. Calendar commitments must follow the actual POP start and Government decisions. The draft sets a six-calendar-month Phase I deadline; do not equate that deadline with 24 weeks or assume acceptance/review latency is zero. Draft schedule.
Workstream and proposed window
Concrete work and owner role
Exit evidence / dependency
Weeks 1–2: baseline and risk probes
Technical lead + audio/ML engineer: requirements matrix, provisional profile, corpus rights, two-voice short sample, protected tiny-book spike, Mac runtime probe
Recorded decisions; encoder/protector/player access requested immediately. A missing codec/protection route blocks a claim of complete DTB success.
Weeks 1–4: design package
Technical lead + security/QA contributors: five specified diagrams and eight specified documents
Submit within 30 calendar days of POP start, regardless of the week labels. Unresolved decisions remain explicit.
Weeks 3–6: common core
Backend/DTB engineer: source store, versioned book schema, EPUB/DOCX/RTF/text import, compiler rules and job state machine
Tiny reference book round trip; no unaccounted source text or broken navigation; stable worker contracts.
Weeks 5–9: narration and review
ML engineer + product/backend engineer: model adapters, pronunciation overrides, issue queue, transcript/audio comparison, versioned approvals
Full text-origin book, both voices, deliberate pronunciation repair with correct downstream invalidation.
Weeks 7–11: commercial path
Backend/DTB + ML engineers: media decoding, transcript alignment, metadata navigation, supplemental speech
Complete commercial-source sample with matching text; audio-only sample with uncertainty surfaced for review.
Weeks 10–13: reliability and evaluation
QA/accessibility lead + all engineers: complete corpus, malformed inputs, crash recovery, bounded retries, independent playback
Measured defect and repair rates, all output checks, corpus coverage, accessible operator flow.
Weeks 13–16: AWS proof and handover
Technical lead + platform/security contributor: approved AWS sandbox, controlled model transfer, integration rehearsal, clean rebuild and demo
Same reference suite passes; compute/storage/review costs recorded; handover rehearsal and test report.
Remaining time before Phase I deadline
Project lead coordinates Government feedback, component delays and corrective work
Preserve contingency explicitly. If required access is late, revise the schedule and report the blocked gate rather than relabel a partial demo as accepted.
Some work overlaps; format/protection proof, approved output rules and lawful voice/model use constrain later release claims. Model experiments and source-preserving import work can continue while external answers are pending.
Concrete implementation shape
Use a modular application before distributing every component. Proposal: Python API and workers, PostgreSQL for state, an accessible browser operator interface, filesystem artifacts on Mac, and S3 artifacts in AWS. A small job queue and worker leases are sufficient initially; production orchestration can use approved AWS Batch/Step Functions after the basic contracts are stable. See architecture and operations.
Profile version, actual delivered files, validator report, checksums and dependency graph
Review and delivery
Human resolves exceptions; release requires current evidence; delivery adapter records receipt
Actor, decision, rationale, approved version, package identifier and acknowledgement
Run local Metal/MLX inference as a native macOS process. Keep database/API containers optional. The Linux CUDA image is a separate build behind the same worker interface, not a claim that a Mac container image has GPU parity. Platform-dependent numerical differences are expected; compare coverage, quality and output validity rather than require identical generated waveforms.
Operator demonstration
Demonstrate a traceable session: import an EPUB, inspect recovered chapters and metadata, choose an approved voice, generate a chapter, inspect a flagged word, correct its pronunciation, audition before/after, rebuild the affected audio/timing/package, run validators, and navigate the protected result on a target player. Then repeat with a commercial audiobook, preserving its narration and adding only the approved generated material.
The UI needs book/job lists, a navigation-tree editor, text/audio issue review with keyboard controls, voice and lexicon settings, an evidence panel and a release action. A release blocked by stale validation should explain which change invalidated the evidence. Separate operator, reviewer/releaser and administrator permissions even if one person holds several roles during the pilot.
Deliverables and staffing
Propose four core contributors: a technical lead/platform engineer, a backend/DTB engineer, a speech/ML engineer and a QA/accessibility engineer. Product/UI design, audio editorial expertise, Library integration/security and procurement/license review are fractional support roles; they are real dependencies, not assumed free capacity. Reserve NLS reviewer and player access in the schedule. With one or two engineers, revise scope or time rather than inherit this estimate.
The day-30 design package should map directly to the SOW's five diagrams: System Content Diagram (wording in the draft), Component Architecture, Data Flow, DTB Packaging/Output Structure, and Security/Access. Its eight documents cover System Architecture, DTB Integration, Text Processing, Audio Output, Metadata/Packaging, Security/Data Protection, Testing/Validation, and Implementation Schedule. The wiki figures explain the proposal; they do not by themselves constitute all these signed-off deliverables. SOW §3.13.1 and §4.1.
Map the technical lead to the Senior Engineer key-personnel responsibility and allocate a separate Project Manager, proposed at 0.25–0.5 FTE during the prototype. The draft identifies both roles in §6.1; actual personnel must satisfy the stated qualifications. The PM owns the decision log, dependency escalation, schedule and status/quality-control reports every 30 calendar days. This management capacity is additional to the four engineering contributors, not hidden in their estimates. SOW §§3.10–3.11, 4.1, 6.1.
End-of-prototype delivery includes a reproducible code release, dependency/model/voice inventory and license notices, configuration/profile exports, corpus manifest, voice comparison, sample DTBs, test results, benchmark raw data, known limitations, backlog, setup/rollback/restore runbooks, a workflow recording and an operator/admin walkthrough. No source books, private recordings or keys go into the public wiki.
Decisions before implementation spending
Resolve the output-profile editions, synthetic-audio mastering interpretation, available encoder/protector and OS support, authorized voices, Library account/region/network access, Section H controls and retention, target players and acceptance owner. Obtain throughput and human-review budgets before selecting production capacity. Decision register.
If an encoder or protection tool is Windows-only, add a controlled Windows worker and a license/deployment review; Mac development does not remove that dependency. If full-book speech quality fails, compare the baseline and challenger or narrow the agreed content class; do not hide failures behind unbounded retries. If a model license cannot meet the ownership/access terms, replace the component before investing in its integration.
Start with Qwen3-TTS-12Hz-1.7B-CustomVoice for narration, and evaluate it against Kokoro-82M. This is a proposed starting point, not a finding that either model meets NLS requirements. Keep model inference local to the developer Mac or the approved AWS environment; the delivered workflow should not depend on a paid speech API. Training or fine-tuning a new model is outside the initial prototype scope.
Primary narration choice
The exact upstream artifact is Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice. Its card lists Apache 2.0, preset voices, multilingual synthesis and style instructions. Choose CustomVoice, rather than Base, for the first prototype because it provides preset speakers without requiring a new voice-cloning workflow. The official examples use CUDA/BF16. Official Qwen implementation.
Start with restrained narration and short sentence-group chunks. Upstream lists Ryan and Aiden as English-native presets, both male. Include other approved English voices in the comparison if the pilot needs a broader selection; multilingual capability alone does not prove equal narration quality for every speaker/language pairing. Treat speaker choice as book-level configuration. Never silently switch engines or voices to repair one sentence.
The official checkpoint tree is approximately 4.52 GB at review time, including associated files. That is download size, not peak inference memory. The “1.7B” name alone cannot determine the total working footprint. Pin the selected revision and record actual memory during the runtime probe.
Small Apache-2.0 model with 24 kHz sample output. Useful for plumbing, announcements and voice comparisons; English voice catalog provides more preset choices. Its size does not establish full-book quality.
MIT, 809M-parameter ASR candidate. Transcribes generated and commercial audio to flag discrepancies. Hallucinations and recognition errors mean its transcript is evidence for review, not truth.
Apache 2.0; aligns supplied text/audio, with a documented maximum of five minutes per input. Use much shorter overlapping windows and reconcile boundaries. Alignment cannot independently prove the supplied words were spoken.
Apache 2.0 instruction model for schema-constrained heading/reading-order suggestions on ambiguous input. Use a proposed 4–8K context initially. It must not rewrite narration, invent catalog facts or generate authoritative final XML.
MIT project with several model variants and an embedded watermark. Compare only if the initial two-engine evaluation leaves a meaningful quality gap; validate exact variant, voice rights, runtime and watermark behavior after encoding.
The prototype needs TTS plus independent audio checking; it does not need every optional model loaded or even integrated on day one. Deterministic parsing remains the first choice for documents with usable structure.
Pinned local inference runtime and the same approved source model; benchmark separately
MLX-Audio is a community implementation. Its support is not a claim that the original CUDA package runs unchanged on Mac. Qualify the conversion on identical passages and on full-book continuity. Metal/MLX uses the Mac's shared memory; a CUDA worker has separate GPU and host memory. Hardware tiers.
For the first quality comparison, keep Qwen TTS at BF16. Quantized TTS is a separate, held-out experiment. Record runtime-specific output changes; do not assume a smaller checkpoint preserves fidelity. Use separate inference environments when dependencies conflict.
Guardrails in the adapter
Each synthesis request carries approved text, language, speaker ID, narration style, pronunciation policy, model and runtime versions, precision and a deterministic job identifier. Save audio duration, sample rate, text hash, elapsed time, memory peak and warnings with the output. Seeded generation improves traceability but does not guarantee identical audio across hardware.
The adapter must reject empty or incomplete output and detect truncation. Kokoro's pipeline implementation has bounded phoneme sequences and truncation-related behavior; enforce complete source coverage before calling inference, not merely by reading logs afterward. Use sentence boundaries and source spans to split material, and keep the distinction between synthesis chunks, navigation units and packaged files.
Pronunciation overrides change a separately recorded spoken representation. Do not feed arbitrary SSML to a model unless its adapter explicitly supports it. Check names, numbers, units and negation even when overall ASR word error rate looks good. Use the test plan for release evidence.
Rights, mastering and handover
Apache/MIT labels are useful starting evidence, not blanket clearance of model provenance, voice rights, dependencies, redistribution or the draft's ownership/non-infringement terms. Maintain a reviewed inventory of exact checkpoint revisions, conversions, voice assets, code licenses and any reference-speaker permission. Voice cloning, automatic character voices and style effects are deferred unless explicitly approved for the content.
Pre-fetch approved artifacts into a controlled cache, verify hashes and prove inference with network access disabled. Continued operation must not require a vendor account, hosted model endpoint or license renewal. The final rights arrangement still needs procurement review against SOW §§2, 3.9, 3.12.
Native synthesis sample rates and token rates are not output-compliance settings. Keep original generated audio, then apply the approved mastering/encoding profile. Upsampling to 44.1 kHz does not settle whether a model satisfies a requirement about initial recording; resolve the mastering question before making a compliance claim.
Proposed starting configuration: an actively cooled Apple Silicon Mac with 64 GB unified memory and a 1 TB SSD, plus a short-lived AWS g6.2xlarge test worker with one NVIDIA L4, 24 GB GPU memory, 8 vCPUs and 32 GiB host RAM. Run model stages sequentially at first. This is a practical provisioning proposal for the model stack, not a measured minimum, throughput guarantee or purchase order. AWS G6 specifications.
What the draft establishes about cloud providers
The supplied draft says to integrate with NLS AWS systems (§3.7), and describes existing S3-based storage (§1). It does not provide a cloud-provider allowlist or explicitly require every workload to execute exclusively in AWS. Because Section H is absent, the permitted account, region, services, data residency, network boundary and security authorization cannot be determined from these attachments alone. Draft SOW §§1, 3.6–3.7, 5.1.
Environment / provider
Evidence of permission
Proposed use
Library-approved AWS environment
AWS integration is explicit; particular compute service/account/region approval is not established
Primary candidate for cloud inference, storage and integration testing. Confirm service availability and GPU quota before booking work.
AWS GovCloud
Not named in the supplied draft
A conditional AWS option only if the Library specifies that boundary; do not assume it is required, sufficient or interchangeable with commercial AWS.
Azure, Google Cloud, Oracle Cloud
No authorization established by the supplied attachments
Do not include as approved runtime choices. Request a written determination before any contract-content processing or deployment there.
RunPod, Lambda, CoreWeave or other GPU hosts
No authorization established by the supplied attachments
Potential technical alternatives only after provider and data-handling approval; excluded from the baseline.
Developer Mac
Proposed engineering environment, not an expressly approved place for Government content
Public-domain, synthetic or explicitly authorized fixtures until content-handling permission is confirmed.
This Cloudflare Pages wiki
Public documentation hosting already requested by the user
Documentation only. Its deployment establishes no permission to host source books, protection keys, private recordings or inference there.
Cloud certifications alone would not establish contractual permission. Keep software-access rights and consumption-based infrastructure costs as separate questions; the draft's non-subscription language should be clarified for the final procurement arrangement.
Mac development tiers
Tier
Proposed hardware
Intended use and tradeoff
Reuse existing hardware
Apple Silicon, 16 GB unified memory; adequate free disk
API/UI/compiler work, Kokoro and short serialized smoke tests. Useful starting machine; not the buying target for sustained Qwen plus QA development.
One model loaded at a time, bounded chunks/context, few competing apps. Verify actual peak memory and swap before accepting it for whole-book runs.
Recommended desktop
Mac Studio, M4 Max, 16-core CPU / 40-core GPU, 64 GB unified memory, 1 TB SSD
Comfortable development headroom for speech experiments and application tooling. Apple lists this memory/storage combination; the recommendation is an engineering judgment, not an Apple TTS benchmark.
Portable alternative
MacBook Pro, M5 Pro with 20-core GPU, 64 GB unified memory, 1 TB SSD; M5 Max optional
Similar memory budget in a portable, actively cooled form. More GPU cores may help, but throughput must be measured on the selected configuration.
Larger workstation
96–128 GB unified memory
Defer unless simultaneous model workers, a larger LLM or training becomes an evidenced need. Extra capacity is not required by this proposed scope.
Apple's published configurations: Mac Studio specifications and M5 Pro/Max MacBook Pro specifications. These establish supported configurations, not current inventory, delivery dates or prices. Confirm availability before purchase. An existing 48 GB machine is also a reasonable evaluation workstation; do not replace it before the memory/throughput probe.
The Mac's memory is shared by GPU, OS and applications; it is not equivalent to the same number of gigabytes of dedicated NVIDIA VRAM. Reserve at least 200 GB initially free for the pilot, plus a backup destination. Use native Metal/MLX workers; validate Linux/CUDA separately. Headphones for careful listening, representative NLS playback hardware and an accessible test browser/screen reader are as important to acceptance as the GPU.
Working-memory budget to validate
Process
Proposed starting allocation
Scheduling policy
Qwen 1.7B BF16 TTS
8–12 GB
Batch 1, bounded sentence groups; measure cache/codec/framework peaks
Kokoro baseline
2–4 GB
Load instead of Qwen for a comparison run
Whisper turbo ASR
3–6 GB
Inspect accepted/generated chunks after synthesis; no required co-residency
Optional forced aligner
2–4 GB
Short windows; reconcile boundaries
Optional 4B LLM at 4-bit, 4–8K context
4–8 GB
Separate structural suggestion job; no whole-book context assumption
OS, database, API, browser and build tools
Reserve roughly 10–16 GB on Mac
Limit container memory and competing apps; watch memory pressure
All allocations above are estimates to test, not vendor minimum specifications. They depend on runtime, precision, input length, batches and conversion. Do not add every row as if all models must reside simultaneously. Conversely, do not use raw checkpoint size as total RAM. Reduce batch/context first, then serialize or unload workers before buying larger hardware. Model evidence and runtime choices.
AWS instances and surrounding resources
Candidate
Published hardware
Prototype role
g6.2xlarge — first choice to test
1 × L4, 24 GB GPU memory; 8 vCPUs, 32 GiB RAM
One sequential inference worker; separate CPU compilation if host memory becomes tight
g6.4xlarge
Same single L4 / 24 GB GPU memory; 16 vCPUs, 64 GiB RAM
More host-side parsing/media/cache headroom; it does not double VRAM
g5.2xlarge — availability fallback
1 × A10G, 24 GB GPU memory; 8 vCPUs, 32 GiB RAM
Benchmark against G6 if approved region/quota or quoted cost makes it preferable
g5.4xlarge
Same single A10G / 24 GB GPU memory; 16 vCPUs, 64 GiB RAM
Host-memory alternative; validate the exact backend independently
Hardware values come from AWS G6 and AWS G5 product tables; GB and GiB retain AWS's published units. Family support does not establish capacity, regional availability or Library approval. Avoid fractional-GPU instances for the first reproducibility baseline. No multi-GPU/H100 reservation is justified by current evidence.
Proposed supporting resources: one 4-vCPU/16-GiB CPU application/compile node (or equivalent approved service), PostgreSQL, a private artifact bucket, a private model cache, encrypted 250 GB scratch/EBS capacity for the GPU node, bounded log retention, and backup storage. The initial CPU/database services may share a development node; isolate them for later production requirements. Use durable storage for checkpoints; instance-local scratch is replaceable. Start with on-demand test windows and shut workers down when idle; test interruption recovery before considering Spot.
Same book contract, different execution environments
↓ the same jobs and evidence requirements run in two environments ↓
Mac development
Native Metal / MLX workerQwen conversion; CPU/MPS baseline; sequential model loading
Local API + state + artifact filesInspect structure, audition changes, track jobs and rebuilds
Public or authorized test corpusProve quality, memory fit and the operator workflow
Approved AWS test environment
Linux CUDA worker · L4 / A10GUpstream Qwen; pinned model/runtime and driver versions
API + PostgreSQL + private S3Durable artifacts, controlled identities and approved interfaces
Authorized corpus and accessRepeat reference tests; measure cost, throughput and recovery
↓ both submit accepted audio and metadata to the approved toolchain ↓
Master → compile → encode → protect → validate → playTool OS support and target players are dependencies; a separate worker may be necessary.
Cloudflare hosts this public plan only. Contract source content, inference and protection keys stay within separately approved boundaries.
Portability means matching behavior and evidence. Mac and NVIDIA runtimes need independent qualification; they need not generate byte-identical waveforms.
Package identical schema versions, text fixtures, voice recipes and output profiles for both platforms. Compare semantic fidelity, timing validity, memory and accepted throughput. Do not require byte-identical stochastic narration across Metal and CUDA. Test the approved encoder/protector on its actual supported OS; if it is Windows-only, budget a separate approved worker rather than assume Linux or macOS support.
Estimate compute and review effort
Change the assumptions to see the capacity effect. Defaults are examples, not observed model speed or an AWS quote. Each worker is a single-GPU instance.
Use nonnegative hours and price; positive processing factors; generation multiplier ≥ 1; utilization 1–100%; and a whole-number worker count. Keep values within the shown limits.
Busy GPU processing
Allocated instance time, total
Compute wall time, approximate
Human review effort, separate
GPU instance subtotal only
No data leaves this page. Changing worker count changes estimated compute duration, not total work. Human review, CPU work, queues and service limits may prevent linear speedup.
The generation multiplier approximates total replacement-audio work, not the maximum retry count. This excludes CPU compilation, optional LLM jobs, repeated full-source analysis, storage, network, engineering and review costs. It assumes enough independent jobs for the requested concurrency.
Measure before sizing production. Substitute observed per-stage factors, actual regeneration and review effort, and a current region-specific quote.
The worksheet uses explicit assumptions; its defaults are illustrative, not benchmarks. It models one GPU per instance and one selected GPU stage at a time. Generated hours include full text narration and supplemental commercial-audio speech. The generation multiplier covers additional synthesis and checking of replacement audio. Commercial original audio is checked once; repeated full-source analysis, optional LLM work, CPU encoding, queue delays and storage/transfer need additional budget. Effective utilization converts busy processing hours to allocated instance-hours; aggregate wall time divides by worker count, while total paid capacity remains unchanged in this simplified model.
Human review is separate and may dominate elapsed delivery. Pricing is deliberately entered by the reader because region, operating system, purchasing arrangement and date change it. Use a current AWS Pricing Calculator estimate for the approved region, and add CPU/DB, EBS/S3, backups, logs, network egress, player access, engineering and human-review costs. Set an account budget and worker time limit before any cloud experiment; no cloud instances are provisioned by this plan.
Storage sizing and buying decision
As an arithmetic example, 44.1 kHz, 24-bit mono PCM uses about 0.476 GB per hour: sample rate × bytes/sample × channels × seconds. Fifty hours with four retained PCM versions is about 95 GB before originals, models, intermediates and backups. This example is not an approved mastering profile. Stereo doubles that quantity; compression and retention change actual needs.
Buy or reserve only after a 30-minute representative narration/analysis run, a long-chapter run and the protected-output spike establish fit. Capture peak memory, sustained swap, per-stage RTF, reviewer effort and tool compatibility. Scale worker count only after a three-book batch and fault recovery pass. Approve production capacity against accepted audio-hours and required turnaround, not model parameter count alone.
These are proposed engineering gates, not Government-approved acceptance thresholds or completed test results. The prototype should produce evidence for both source routes, full-book speech fidelity, protected playback, repair behavior and operational handover. The draft requires quality checks and functioning Library playback; it does not supply numerical TTS quality or throughput targets. SOW §§3.5, 3.11, 3.13.
Corpus with a reason for each item
Use public-domain, created-for-test or explicitly authorized content on developer machines. Record permissions, source hashes, language, duration/word count, expected navigation and known difficult passages. Government-provided materials require the approved handling environment. Proposed corpus: six complete books, about 30–50 source/output audio-hours in total, plus a short adversarial suite. The range is a collection target, not a statement about available books or a statistical guarantee.
Corpus component
Coverage
Human reference work
Four text-origin books
One each in EPUB, DOCX, RTF and plain text; include prose fiction, nonfiction with notes, numeric/abbreviation-heavy material and weakly structured text
Establish authoritative reading order, metadata, headings and approved spoken forms; distribute difficult features across books.
Two commercial-origin books
One with matching text/metadata; one audio-only with limited metadata
Reference chapter boundaries and announcements; explicitly mark unknown page information.
Format fixtures
Short valid WAV, MP3, OPUS, FLAC plus all four text formats; corrupt, encrypted and misleading-extension examples
Known success/rejection result for every fixture. Avoid lossy re-encoding merely to enlarge the corpus.
Difficult-passage set
At least 100 approved excerpts with names, numbers, dates, abbreviations, notes, repeated words, punctuation, long sentences and language switches
Reviewer-approved pronunciation and reading policy; hold out 20 excerpts from tuning.
Tiny protected reference book
Three chapters, nested navigation, announcement, a note and one known page reference
Exact expected metadata/navigation plus target-player checks; use an NLS-approved reference package where available.
Compare two approved voices and two engines on the same difficult-passage set. Select one recipe before the expensive complete-book run, retain the holdout set, and use a second reviewer to adjudicate disagreements. A six-book corpus demonstrates coverage; it cannot establish a universal error rate for all genres or languages.
Exit gates
Gate
Proposed pass condition
Saved evidence
Source fidelity
Every source span accounted for; every omitted, expanded or normalized spoken passage has an explicit rule or reviewed decision
Source-to-spoken-text map, coverage report, normalization diff; ordinary punctuation transformations recorded by rule
Book structure
All required metadata and expected navigation targets are correct in the fixtures; no broken references or unsupported invented page numbers
XML/schema and profile validation, expected-vs-actual navigation comparison
Speech quality
Full human listening pass of final encoded/decoded release audio for the six-book corpus, including supplemental announcements; zero known unresolved omissions, hallucinated passages, repetitions or material meaning/pronunciation defects in release candidates
Timestamped defect ledger, repairs and reviewer sign-off; report residual uncertainty and minor defects separately
Voice suitability
Blind comparison of clarity, stability, fatigue and joins; NLS reviewers agree voice acceptance threshold before scoring
Per-voice distributions, sample counts, reviewer disagreement and approval; do not claim a generic MOS score proves acceptance
Timing and mastering
Measured outputs satisfy the approved profile's clip boundaries, levels, sample rate and encoding rules
Automated measurements plus spot-checked waveform boundaries; timing explanation
Encoding and protection
Actual AMR-WB+ plus PDTB/NLS protection using approved tools; output decodes/plays through the authorized path
Tool versions, protected package hashes, validator results, authorized-key playback; never include keys in reports
Playback
Every navigation entry in the tiny book and all top-level/representative nested entries in long books work on the agreed NLS hardware/application matrix
Device/application versions, exact actions and observed results; browser audio preview alone is insufficient
Repair correctness
Pronunciation, paragraph-length and heading-label changes rebuild every dependent artifact and preserve unaffected accepted chunks
Dependency comparison, new timeline/checksums, old validation invalidated, post-rebuild playback
Reliability
Three-book queue survives a worker kill, duplicate request, interrupted upload and restart without lost decisions or duplicate release
Job-event trace, recovered artifacts, deliberate fault results; proposed retry cap: two regeneration attempts after the first
Operator accessibility
Complete import/review/repair/release workflow by keyboard and agreed screen-reader combinations; focus, labels, issue announcements and audio controls usable
Manual task test, assistive-technology versions and findings; formal standards baseline agreed with the Library
Security and handover
Unauthorized release/content access denied; logs omit source content/secrets; unapproved egress denied; agreed retention/deletion policy verified across originals, intermediates, caches and backups; backup restores a reviewed book; clean installation reproduces the reference workflow
Access, retention/deletion and restore tests, dependency inventory, offline model-loading test and another operator's setup log
“No known unresolved defect” means the agreed review found and resolved defects; it does not prove the model can never produce an error. Full listening is deliberately expensive for this small pilot so the team can measure what an efficient production review policy would miss. ASR agreement alone must never grant release.
Benchmark procedure
Freeze model, voice, text normalization, chunking, precision, runtime, driver/OS and machine configuration. Record exact model revision hashes; do not compare mutable latest labels.
Run cold-load and warm inference separately. Use fixed short passages first, then one complete chapter, then whole books. Record output duration and errors along with elapsed time.
Measure TTS, ASR/alignment, CPU mastering/encoding/protection, queue wait, transfer and human review separately. Report p50/p95 chunk latency, peak memory, swap, failure and repair rates; include sample counts.
Repeat representative runs, including long inputs and interrupted jobs. Count failed generations, discarded audio and repair attempts in cost, rather than measuring only accepted best runs.
Repeat the same corpus on Mac and the chosen AWS instance. A faster machine is useful only if accepted output and repair effort remain acceptable. Qualify quantized and converted model artifacts separately.
RTF (real-time factor) = processing seconds ÷ resulting audio seconds for a stated stage. RTF 0.5 means that stage takes half as long as its audio; RTF 2 takes twice as long. Do not use time-to-first-audio or streaming latency as whole-book throughput. The hardware worksheet is a planning tool until measured inputs replace its examples.
Report accepted audio-hours per paid instance-hour, cost per accepted audio-hour, and reviewer minutes per accepted audio-hour. Suggested feasibility target for discussion: combined automated processing at or below the audio duration on the selected cloud worker, with no sustained memory pressure. This is a proposed test target, not a measured result or an NLS requirement; failure should trigger a sizing/quality tradeoff, not suppression of data.
Traceability and report format
Each requirement row should contain an ID, exact SOW/specification edition and section, implementation version, test fixture, expected result, observed result, evidence URI, defect ID and acceptance authority. Status values should distinguish not tested, passed, failed, blocked and not applicable with rationale. A missing player or protector is blocked, not passed.
The final report includes corpus coverage, raw measurements, voice comparisons, issue severity definitions, resolved and open defects, security/environment exceptions, playback evidence and the next-phase recommendation. Record both successes and excluded material. Acceptance should be attributed to the authorized reviewer; the engineering team's internal gates do not replace Government acceptance.
Demonstration agenda
Allow about 60–90 minutes as a planning assumption: inspect the corpus/profile and rights; create a short text-origin book; audition two voices; review and repair a flagged passage; inspect regenerated navigation and hashes; show the commercial-audio path; navigate protected outputs on target hardware; inject/recover one worker failure; and review benchmark and handover evidence. Pre-generated complete books complement the live short example so the demonstration does not mistake a polished excerpt for long-form validation.
The draft statement of work describes software for producing NLS digital talking books through two input routes. The engineering proposal is to share the downstream production system between those routes.
The draft background describes approximately 3,000 DTBs produced annually in the existing operation. That is context, not a numerical throughput acceptance target for the new system. SOW §1.
The two required input routes
Route
Core work
Important boundary
Existing commercial audiobook
Ingest audio and metadata, identify divisions, create required supplemental speech, and construct the DTB.
Proposed handling: preserve the supplied narration through approved mastering/conversion. A transcript supports analysis rather than replacing the recording.
Electronic text
Ingest EPUB, Word, RTF, or plain text; recover structure; generate speech; construct the DTB.
Preserve the source and document any change to its spoken form.
The commercial route explicitly accepts WAV, MP3, OPUS and FLAC, and requires AMR-WB+ conversion with MP3 encoding optional. The text route requires WAV and AMR-WB+ outputs, with MP3 encoding optional. These routes and formats come from SOW §§3.1–3.2. Scanned-book OCR is not an explicit core input in the draft. Treat it as a potential extension that needs scope agreement. A raw scan also creates recognition errors that the current text-first workflow would otherwise avoid.
Requirements and proposed evidence
The first two columns summarize the draft; the last is a proposed way to demonstrate completion, not a Government-approved acceptance checklist.
Area
Draft requirement or concern
Proposed evidence
Outputs
NLS-compatible DTBs, navigation, metadata, required audio and related files; §§3.1–3.2
A reference book from each input route, validator results, and playback observations.
Processing
Batch operation and automation with useful operator control
Repeatable batch run, failed-job recovery, and an exception-review demonstration.
Integration
Work within the Library's AWS environment and interfaces
Deployment in the approved environment and a traced delivery to the agreed destination.
Security
Federal security requirements referenced through Section H; §3.6
An agreed controls matrix after the missing section is supplied.
Compatibility and testing
Library playback environment; §§3.11, 3.13
Test records for the agreed hardware, applications, and protected packages.
Ownership and continuity
Non-subscription delivery and Library ownership; §§2, 3.12
Deliverable inventory, reproducible deployment, and dependency/voice-license records.
Documentation and support
Technical documentation, training/support and warranty obligations
Manuals, deployment runbook, handover exercise, support process and warranty terms.
Use the full SOW as the authority for exact wording; this table is a reading aid. The milestones page captures the concrete schedule terms.
Why local inference is the proposed default
The ownership and non-subscription language favors a system that can keep running without a required commercial speech API. The proposal is to deliver application code, approved model artifacts, configurations, and deployment machinery into the Library's environment.
That is an architectural interpretation, not a finding that every hosted service is forbidden or that open weights automatically satisfy the contract. AWS infrastructure still costs money. Third-party code, model weights, voice recordings, and commercial components each need their own rights assessment. Library ownership of delivered work does not erase third-party licenses.
Scope boundaries to settle
Agree which languages, book types, complex tables, equations, poetry, footnotes, and mixed-language passages are included in acceptance. Agree whether page navigation is required only when authoritative page references exist. Decide whether “various multimedia types” means an initial audio-plus-navigation product or additional full-text outputs. See open decisions.
A practical traceability record
For each accepted requirement, record its source section and edition, implementation component, test procedure, evidence artifact, and acceptance status. Keep this separate from model confidence scores. A model saying that a book looks correct is not proof of contractual acceptance.
Example proposal: the requirement for a valid navigation target maps to the compiler's reference checker, a deliberate broken-target test, and a playback test that reaches the expected chapter. This connects a rule to observable behavior.
The supplied materials do not identify one unambiguous standards baseline. The draft repeatedly references Z39.86-2002, while the supplied ANSI/NISO attachment is the 2005 edition reaffirmed in 2012. Record and resolve this conflict before building an exporter.
Editions in the source package
Source
Identified edition
How to use it
Draft SOW
ANSI/NISO Z39.86-2002; applicable NLS specifications, with current-version language
Starting statement of requested work; request an explicit precedence ruling.
Comparison reference; not an automatic replacement for J1.
The public 1203:2025 specification still references the 2002 standard and removes the older 250-file limit. The attached edition and the public edition therefore need separate profiles until NLS confirms which applies. Public construction specification.
Why the edition changes code
An edition can change element names, allowed structures, metadata, counts, or validation behavior. For example, the 2005 NCX has pageList and pageTarget; the 2002 standard uses navList and navTarget for page navigation. A generic DAISY exporter can produce valid output for its own target and still miss the required NLS target. Compare the 2002 standard with the attached 2005 standard.
The SOW's reference to multiple multimedia types also needs clarification against J1's audioNCX metadata requirement (§5.3). The proposal is to deliver the approved audio-and-navigation profile first while preserving internal text structure for possible future outputs.
Proposed versioned profile
Treat an output profile as configuration with an immutable version, not scattered constants throughout the code. It should identify:
Governing standards and exact schema/DTD/entity files, with hashes.
Required metadata, navigation classes, announcements, and supported multimedia types.
Master-audio and compressed-audio parameters, timing rules, and container constraints.
File names, file-size/count limits, checksum format, protection rules, and ZIP behavior.
Approved validator versions and target playback systems.
Documented NLS clarifications or authorized deviations, including their scope.
Every build records its profile ID. A change in profile creates a new build lineage so that later reviewers can explain why two otherwise similar books differ.
A profile selection process
Send NLS a short matrix of conflicting requirements, with source sections.
Ask for written confirmation of editions, precedence, output types, and supplied tools.
Turn the approved rules into a machine-readable profile and human-readable checklist.
Compile a small reference book and test it with the approved tools and players.
Freeze that profile for the pilot; handle later amendments as explicit version changes.
These are proposed implementation steps. The wiki does not establish legal or contractual precedence between conflicting documents.
Current references are separate from the baseline
The public mastering, protection, and delivery pages should be checked alongside the approved profile. Preserve dated copies or hashes when implementation begins; a live link may change after this wiki is written.
The proposed system has two front ends and one production core. Electronic text needs parsing and narration. Existing audiobooks need structure recovery and supplemental speech. Both converge on the same book representation, audio processing, compiler, validation, and delivery stages.
How the parts connect
Both input routes preserve their source evidence and feed the same book model, compiler and acceptance process.
Proposed architecture. The input route changes preparation and speech generation; both routes use the same controlled construction and acceptance process. Failed checks return to the affected stage for repair.
Who does what
Work
Proposed approach
Reason
Recognize uncertain structure and ambiguous text
Deterministic parsing first, model-assisted proposals when useful
Source evidence stays available and uncertainty can be reviewed.
Generate narration and announcements
Approved, replaceable speech models
Voices and model families can be evaluated without rewriting the compiler.
Maintain IDs, source coverage, references and timing
Typed data structures and conventional code
These properties need exact, repeatable checks.
Construct XML, encode, protect and package
Versioned deterministic compiler plus approved media tools
Compatibility depends on exact rules and artifacts.
Accept a book
Agreed validators, playback tests and authorized review
Passing a model check is only one source of evidence.
The main modules
Ingestion stores the original files and parses format-specific structure. The book service manages immutable versions, source spans, metadata claims and approvals. Speech and analysis workers synthesize or inspect bounded segments. The compiler uses accepted content and a frozen timeline to emit the required files. Independent validators report defects. The review interface lets an operator resolve exceptions and trigger dependent rebuilds.
The delivery adapter publishes only accepted, complete outputs through the agreed NLS interface. It should not infer an interface from the fact that the environment uses AWS.
Separate three kinds of segmentation
A speech chunk is a unit that can be regenerated. A navigation element is a destination a reader can select. A delivered audio file is a packaging and playback unit. They do not need to have matching boundaries.
For illustration, one chapter could contain 80 synthesis chunks, 12 navigable sections, and one delivered chapter audio file. Those numbers are explanatory, not requirements. Keeping the concepts separate enables local repairs while respecting file limits and meaningful navigation.
Illustrative chapter · counts from the example below · not a timing scale
Three kinds of segmentation solve different problems. A correction can replace a small speech chunk while navigation and delivery remain organized at larger scales. The counts are illustrative, not NLS limits.
A book behaves like an incremental build
A pronunciation correction can affect several speech chunks. Regenerate those chunks, reassemble the affected masters, recalculate the dependent timeline, rebuild navigation/XML, encode and protect the changed content, then update checksums and packaging. Unaffected accepted synthesis can stay cached.
This is a dependency graph: each output records the exact inputs from which it was made. If an input changes, downstream evidence becomes stale. The build system should never keep a green validation result for an old version of a file.
Boundaries for AI assistance
Model outputs are proposals or generated media until verified. They must not silently delete source text, invent page numbers, substitute a narrator, write arbitrary final XML, or mark their own output accepted. Books and metadata are untrusted content, including instructions that appear inside them.
A useful model response is structured: a proposed chapter boundary, the source span supporting it, alternative interpretations, and confidence evidence. The compiler consumes an accepted decision, not a free-form conversation.
Small services before many services
The module boundaries can initially live in one application with separate worker processes. Splitting every module into a network service is not necessary to prove the approach. Establish stable data contracts, job IDs and reproducible artifacts first; deploy independently where isolation or scale justifies it.
The canonical book model is the system's shared record of what the book contains and how each part becomes audio and navigation. “Canonical” means that downstream stages refer to the same structured representation; it does not mean that an AI transcript replaces the original source.
Proposed invariant: every source passage is represented in the output or has an explicit, reviewable reason for exclusion. This is the strongest protection against quiet omission, duplication and paraphrasing.
Follow one passage through the system
Proposed data model · trace one passage through a build
Source passageOriginal text + stable location$1,205
→
Approved spoken text“one thousand two hundred five dollars”
First arrow: retain the normalization rule and approval. Second arrow: retain the model, voice and lexicon versions.
Source span + stable node IDPreserves hierarchy and reading order
→
Structure recordHeading, paragraph, note, table or page marker; parent + order
Structure + accepted audioResolves logical destinations into media
→
Two navigation referencesMain narration destination and separate spoken heading-label range
Approved text + observed audioASR is evidence of what a recognizer heard
→
Quality recordArtifact version, check, evidence, defect and resolution; ASR stays separate from approved text
Exact input versions + dependent checks↓
Auditable output buildProfile and tool versions, dependency records and package hashes identify exactly what was delivered.
One passage, connected records. This proposed model preserves the source while linking each transformation, media asset and approval to an exact version. The currency example illustrates normalization; it is not an NLS-prescribed phrase.
What to retain
Record
Essential information
Why it matters
Source artifact
Original bytes, hash, format, ingestion version
Reproduce and audit a conversion.
Source span
File and stable structural location; original text or audio range
Locate the evidence behind every transformation.
Structure node
Heading, paragraph, note, list, table, page marker; parent and order
Preserve reading order and navigation semantics.
Text transformation
Original text, approved spoken form, rule, rationale, reviewer
Distinguish faithful normalization from accidental rewriting.
This schema is a proposal. Its final field names and storage technology are implementation choices.
Preserve three forms of text
Store original written text, approved spoken text, and observed ASR text separately. Each answers a different question: what did the source say, what should be spoken, and what did a recognizer hear?
Example:
Original text: $1,205
Spoken text: one thousand two hundred five dollars
Transformation: currency/USD
Source location: chapter 4, paragraph 12, characters 18–24
ASR evidence: populated only after examining the generated audio
A normalization rule should never convert “fifty” and “fifteen” into the same comparison token. Preserve high-consequence details such as negation, units, dates and proper names.
Use stable IDs and immutable versions
Give structural nodes stable identities across edits when they represent the same passage. Give source snapshots and generated assets immutable version IDs or content hashes. A node can point to a new approved audio asset without overwriting the old one.
A timestamp alone is an inadequate identity: earlier edits can shift every later timestamp. Store the logical destination and the resolved audio position, then derive the latter for each completed build.
Track two independent audio references
A navigation item points to a destination in the main narration and may also reference a spoken label in a separate headings file. Record both asset identities and timing ranges explicitly. Confusing the two timelines can produce a correct-sounding label that jumps to the wrong place.
Explore which work changes
Proposed dependency model · explore a change
A lexicon correction regenerates affected speech. Downstream timing and packages follow the changed audio; unrelated accepted synthesis can be reused.
Rebuild or rerunReuseDepends on scope
1
Original source + structureReuseRetain the source snapshot and stable passage identities.
↓ source spans + approved spoken forms
2
Main narration chunksRebuild affectedRegenerate passages using the corrected lexicon entry; reuse unrelated accepted chunks.
Spoken heading labelsDepends on scopeRegenerate only labels that use the corrected entry.
↓ accepted media versions from both branches
3
Assembled audio + final timingRebuild affectedReassemble affected masters and calculate final sample positions.
↓ resolved destinations + clip ranges
4
Navigation + XMLRebuild dependentRebuild dependent SMIL/NCX references and durations from the new timeline.
↓ compiled content + output profile
5
Encode, protect + packageRebuild changedProcess changed media and update dependent protection, checksums and packages.
↓ exact new artifact hashes
6
Validation evidenceRerun dependentValidate the changed dependencies and completed build; prior approvals remain tied to prior hashes.
Scope matters: a surname can have different pronunciations in different contexts. Record which occurrences the accepted correction covers.
A local edit can reach the final package. This simplified dependency order illustrates the proposed rebuild strategy. Narration and heading labels have separate assets and timing references. The approved profile determines whether a heading clip is copied or synthesized. The implementation must follow the actual dependency graph.
Worked repair example
An operator corrects a surname in the book lexicon.
Dependency lookup finds the passages whose spoken forms use that entry.
Those passages receive new speech recipes and audio versions.
The affected assembled masters and their final sample positions change.
Dependent SMIL/NCX timing, durations, checksums, protection and packages are rebuilt.
Validation runs against the new artifact hashes; the prior build remains auditable.
The system should record who accepted the correction, why it was made, and whether all occurrences should change. A name may intentionally have different pronunciations in different contexts.
Define cache identity carefully
Include source/spoken text, model and voice artifacts, normalization and lexicon versions, generation parameters, and relevant processing versions. Store the accepted audio itself. A random seed cannot guarantee byte-identical future synthesis across different hardware or software versions.
Ingestion should retain source evidence before attempting to improve it. The workflow differs by input format, but every importer should produce source-linked content and proposed navigation for the common book model.
Visual styling that disguises semantic structure; broken reading order.
Word
Heading styles, outline levels, lists, tables, footnotes and endnotes
Direct formatting that looks like a heading but has no heading style.
RTF or legacy Word
Convert in an isolated process while keeping the original and a conversion record
Lost notes, tables, symbols, or outline information.
Plain text
Delimiters and deterministic heuristics, followed by model-assisted classification
Missing hierarchy and ambiguous headings; uncertain reading order.
The format list comes from the SOW §§3.1–3.2; the importer design is proposed. Do not flatten a structured source into plain text before extracting its relationships.
Main content determines navigation
J3 §3.1 directs navigation work to the main content rather than simply copying the table of contents. A table of contents may omit subheadings or describe them differently. Inspect the actual text and preserve evidence for every proposed division.
J3 also separates hierarchical navigation from nonhierarchical structures such as notes and segments. A note is not necessarily another chapter. The correct distinction affects playback behavior and must survive ingestion into the compiler.
A model can propose that a passage is a section heading, but the proposal should include its source location and supporting evidence. Printed page numbers should come from an authoritative page map or source, not model invention.
Commercial recordings
Use evidence in this proposed order: publisher metadata and chapter boundaries, matching electronic text when available, an ASR transcript, semantic classification, and precise audio alignment. Record conflicts rather than silently choosing whichever source arrived last.
For example, a model can distinguish a spoken chapter announcement from dialogue mentioning “chapter three.” The final destination still needs alignment and a waveform-level boundary check.
Keep the supplied narration through approved mastering and conversion. The ASR transcript is an analysis aid that can contain errors. Synthesize only the additional speech the approved output requires. Availability of a narrator's recording is not permission to clone the narrator's voice.
A recording cannot reliably reveal an unspoken print heading or an unannounced print page number. Mark these as unavailable, and agree how that input class will be accepted.
Written text to spoken text
Normalization expands abbreviations, dates, numbers and symbols while retaining original source spans. The sentence “Dr. Smith paid $1,205 on 03/04/2026” needs context: “Dr.” may require disambiguation, the currency needs expansion, and the date convention is ambiguous.
Use deterministic rules for well-defined cases and a reviewable proposal for contextual cases. A book-level pronunciation lexicon keeps repeated terms consistent. NVIDIA NeMo text processing is a candidate foundation for text normalization, not a guarantee that the whole book is correctly interpreted.
Model adapters and source safety
Each speech adapter should declare which controls it supports: SSML, phonemes, pronunciation dictionaries, language selection, reference audio, and style instructions. Unsupported controls must fail clearly or use an agreed conversion; they should not be silently ignored.
Treat model-specific control strings in source text as content. Escape or otherwise neutralize them before a model can interpret them as a voice or performance instruction. Retain the literal source and record the transformation.
What an importer should report
Proposed outputs include an ordered content tree, source coverage map, metadata conflicts, lost-format warnings, navigation proposals and unresolved passages. A successful file parse alone is insufficient: compare recovered content and reading order with the source before moving into speech production.
Use a replaceable text-to-speech (TTS) interface and evaluate several local models on the same material. The names below are candidates from the original analysis. None has been demonstrated here to satisfy NLS acceptance requirements.
Prototype model selection
The prototype model plan selects Qwen3-TTS 1.7B CustomVoice as the initial narration candidate, with Kokoro as a comparison baseline. It narrows the original broad shortlist into explicit Mac/CUDA runtimes, supporting models and evaluation conditions.
Model-card licensing requires attention, including separate commercial licensing
Conditional evaluation candidate after terms are resolved.
License labels do not establish rights to every associated voice, dataset, dependency or deployment use. Pin the exact artifacts and record applicable terms in the handover inventory. This wiki's model descriptions are not a legal license clearance.
A book level narration plan
Before synthesis, select an approved voice, pronunciation lexicon, language handling, and restrained pacing/style settings. Preserve these across chunks. Faithful reading is the default: selectable voices do not imply automatic character acting, added laughs, paraphrasing or sound effects.
Test paragraph or sentence-group chunks. The original analysis suggests experimenting with 20–60 seconds of speech per chunk; this is a tuning range, not a requirement or measured optimum. Preserve sentence boundaries, and compare voice continuity, pronunciation, joins, and correction cost at different lengths.
Bounded generation and repair
Synthesize a chunk from approved spoken text and a recorded recipe.
Transcribe with an independent recognizer and compare meaningful content.
Check alignment, clipping, silence, level and joins.
Repair a specific discrepancy through a lexicon, text, boundary or generation change.
Escalate after an agreed attempt limit rather than retry indefinitely.
Do not silently switch narrators for a difficult sentence. An engine fallback must preserve the approved voice or trigger a deliberate review and potentially a larger regeneration.
ASR and alignment candidates
Qwen3 ASR provides transcription and a separate forced aligner. WhisperX combines transcription with alignment and documents limitations. Long books need bounded jobs with reconciled overlap, rather than a single whole-book alignment call.
ASR asks what a recognizer hears. Forced alignment asks where expected text fits the audio. These are different checks, and neither establishes perfect narration alone. See the validation page.
A useful evaluation corpus
Include ordinary prose and long chapters alongside names, dates, quantities, negation, abbreviations, dialogue, notes, lists, tables, mixed-language passages, and repeated terms. Add passages from each supported input class. Keep tuning material separate from a held-out acceptance set.
Measure complete books as well as excerpts. Report unresolved defects per accepted audio hour, review minutes, repair attempts, throughput, and consistency across joins. Do not infer whole-book quality from a 30-second demo or time to first audio.
Voice governance
Keep approved voice assets, consent or usage rights, version, language scope and allowed transformations with the production recipe. A change to the voice reference can affect every generated passage and therefore should invalidate the appropriate cache and quality evidence.
The compiler converts an accepted book model and finalized audio into the files required by an approved NLS output profile. It should serialize typed data and enforce explicit rules. Generating final XML directly with an LLM would make exact compliance harder to control.
What the main files do
File or format
Role
OPF
Describes metadata, the package contents and reading/playback organization.
NCX
Defines navigable structures, labels and destinations.
SMIL
Connects ordered playback structures to particular media clips.
WAV masters
Accepted uncompressed source audio used as the timing reference.
AMR-WB+ in 3GP
Compressed audio with NLS-specific codec and container constraints.
Resource files
Supply approved alternative navigation-class labels when needed.
dtb.md5
XML checksum document for the DTB files, excluding itself.
DTD and entity files
Local dependencies needed to interpret and validate the XML.
These descriptions summarize J1 §§3.1–3.10. Exact applicability depends on the selected profile.
Freeze audio before final timing
Proposal: track positions as integer samples, complete approved mastering and assembly, then freeze the accepted WAV timeline. Derive clip timestamps from that final timeline, not from rough ASR estimates or an earlier synthesis chunk.
In attached J1 §3.2.5, timestamps remain relative to the accepted WAV; codec offset compensation is assigned to the player. Do not silently shift XML to encoded-audio timing or compensate twice. J1 pp10–11.
Timing windows depend on clip type
Hatched band = permitted location of the clip boundary
Main SMIL narration + heading labels
Both clip boundaries must fall in silence.
Before speech starts
−200 ms0
Start: 80–120 ms before
Speech
duration schematic
After speech ends
0+500 ms
End: 150–300 ms after
Alternative-class resource audio
Use the resource-specific window; its allowed range differs.
Before speech starts
−200 ms0
Start: 100–200 ms before
Speech
duration schematic
After speech ends
0+500 ms
End: 150–500 ms after
Read from the speech edges outward. Each zero marks the adjacent speech boundary. Lead-in and tail use separate scales; the speech block has no time scale. Select one endpoint within each permitted band. These are boundary windows, not required amounts of inserted silence. If a valid boundary is unavailable, flag the case or approve an audio edit and rebuild timing.
The first two rows also require the boundaries to lie in silence. Both clip endpoints must be present. A shared global “navigation padding” constant would miss the resource exception.
Use alignment plus local waveform analysis to identify the actual speech boundary. If the required margin is unavailable, flag the case or make an approved audio edit and rebuild dependent timing. Choosing the middle of a permitted window is a proposed default, not an NLS rule.
Main destinations and spoken labels
A navigation destination resolves into primary audio through SMIL. Its spoken label may refer to a different range in the headings audio. Keep both references explicit. Labels must represent the heading, not the whole section. J1 §§3.4.3.4–3.4.3.5, pp13–14; J3 §3.1.2, p6.
J3 requires at least two level-one navPoints: the first marks the opening title/author with class title/author, and the last marks the closing announcement with class close. J3 §3.1.1, p5.
Notes, note references, sidebars, pages, line numbers and segments need their own semantics. J3 requires segment navigation for entries in specified lists, and J1 says not to create navLists for segments. Treating every navigation object as a chapter loses these distinctions. J3 §§3.2.1–3.2.4, pp6–7; J1 §3.4.4, p14.
Exact rules from the attached construction edition
Rule in J1
Compiler consequence
No more than 250 files; §3.1.3
Count all required files under that edition. The newer public edition differs.
Split SMIL above 100 KiB; at most 50 SMIL files; §3.3.12
Measure actual UTF-8 serialized bytes; 100 KiB is 102,400 bytes. Follow the stated escalation rule if limits are exceeded.
At most 5,000 navPoints; contact monitor when exceeding 1,000; §3.4.3.6
Distinguish a hard limit from a consultation trigger.
Lowercase filenames and prescribed production IDs/suffixes; §3.1
Generate names from the profile and NLS-supplied identity.
XML diskcheck checksum structure; §3.9
Ordinary md5sum text is insufficient. Include each file except the checksum itself; do not list the checksum in the OPF manifest.
Include referenced DTDs and entities; UTF-8 XML; §3.10
Package dependencies and validate against an approved local catalog.
Duration follows SMIL playback, including specified custom-test content; §5.3(v)
Derive duration from the compiled playback structure, within the stated ±1 second accuracy.
See J1 pp9–20. Do not apply the attached 250-file limit to a newer profile by accident. Permit only approved local XML dependencies; do not enable arbitrary network/entity resolution for book content.
Encoding is an early interoperability risk
AMR-WB+ is not ordinary AMR-WB. Attached J1 specifies mono constant bitrate, frame type 23 and ISF index 8, plus particular 3GP metadata and field constraints. A file extension alone proves none of those properties. J1 §3.2.3, p10.
Hindenburg's support article describes an NLS-supplied encoder; APH's revision history records NLS protection/validation integration. These are historical precedents. Obtain current components, terms, supported operating systems, automation interfaces and cloud-deployment approval before relying on them.
Protection is part of the book
PDTB protection involves protected content and associated authorization/key material. S3 storage encryption protects stored objects but does not create a protected talking book that an NLS player can authorize. See the DAISY protection specification and request the governing NLS 1205 profile and tooling.
Packaging and final submission
The public 1206:2025 delivery specification describes an outer ZIP with protected/unprotected DTB archives and WAV masters. It specifies uncompressed entries and local-header CRC/size values with ZIP bit 3 clear. A seekable staging file is a proposed way to control those headers, followed by archive inspection and extracted-file checks. Delivery specification §§3–3.2.
Its §4 names the NLS Transfer Portal. Internal S3 staging is a separate architectural choice; confirm the project-specific final interface. Keep QA logs outside the delivery package unless that interface expressly permits them.
Reuse existing DAISY components selectively
DAISY Pipeline's DTBook-to-DAISY3 script accepts a 2005-3 DTBook input and offers speech-related conversion options. It is a useful candidate for selected transformations. That documentation does not establish an NLS 2002, AMR-WB+, protected export path. Keep the NLS profile and acceptance tests explicit even when reusing its code.
Automation succeeds when defects are detectable and repairs are precise. A generated file is only an intermediate result until content, audio, navigation, format and playback checks agree with the accepted requirements.
The following layered QA design is proposed. The draft requires quality and compatibility testing but does not provide numerical narration-error or review-effort thresholds. Those must be agreed before the pilot is judged. SOW §§3.5, 3.11, Phase II.
From a candidate to an accepted release
Generate a candidateBook model, narration, navigation and package
↓ candidate + artifact hashes
Check six layersSource · speech · acoustics · navigation · format · playbackModel checks are evidence, not acceptance.
↙ issues found
Bounded repair loop
Triage the findingIdentify the cause and the smallest valid repair.
↓ repair is clear; budget remains
Repair affected artifactsRegenerate, correct or make an approved edit.
↶ Return to the six checksRecheck every result invalidated by the repair, including dependent timing.
Ambiguous issue or retry budget spent?→ Escalate to an authorized reviewer. Record the decision and evidence.
↘ no unresolved blockers
Separate release decision
Final acceptance gateApply the agreed content, audio, navigation and delivery criteria.Review evidence independently and verify the agreed player/app matrix.Document any authorized exceptions.
↓ accepted against agreed criteria
Accepted release
← Rejected: return to triage; preserve the failed build and findings.
Check results belong to a specific artifact hash and output profile. A repair creates a new candidate; earlier passing results may no longer apply.
Proposed control flow. The retry budget, acceptance criteria and review authority must be agreed. The draft calls for quality and compatibility testing; the numerical acceptance thresholds still need agreement. Draft SOW §§3.5, 3.11 and Phase II.
Six layers of correctness
Layer
Example defect
Proposed evidence
Source fidelity
A paragraph is omitted or reordered
Source-span coverage and ordering checks against the imported book.
Speech fidelity
A number changes or a word is repeated
Independent transcription, protected-token comparison and targeted listening.
Acoustic quality
Clipping, an abrupt join, long silence or voice drift
Signal checks plus excerpt and sustained-listening review.
Navigation
A label announces one chapter and lands in another
Reference resolution, alignment and interaction tests.
Format
Broken XML references, wrong codec, bad checksum or ZIP headers
Schema, semantic, media and package inspectors.
Playback
Protection, skip, resume or speed behavior fails on a target player
Recorded test matrix using actual agreed players/apps.
ASR evidence is not ground truth
An ASR mismatch can arise from a TTS mistake, a recognizer mistake, or a legitimate difference between written and spoken forms. Forced alignment fits expected text to an audio signal; it cannot independently prove every expected word was spoken.
Combine independent transcription, alignment, acoustic evidence and targeted checks for names, quantities, units, dates and negation. For high-consequence discrepancies, use another recognizer or a human listener. Agreement between models increases evidence but does not guarantee correctness. WhisperX's documented limitations illustrate missing timestamps and other alignment problems.
A review screen should show enough evidence
For each issue, present the source passage, approved spoken text, recognizer output, short playable excerpt, timeline/waveform, reason for the flag, affected build and repair history. Offer concrete actions: correct pronunciation, edit a proposed boundary, regenerate a passage, or accept a documented exception.
An exception should retain rationale and authorized reviewer identity. Acceptance of one passage should not automatically suppress the same error class throughout a book.
Defect states and bounded retries
A proposed workflow is detected → triaged → repair attempted → rechecked → accepted or escalated. Set an agreed retry budget per issue. Distinguish processing failures, which may be safely retried, from content ambiguities, which often need a decision.
Retest everything invalidated by a repair. Changing a chunk's duration can invalidate later clip times even when the later audio remains unchanged. Tie every check result to the exact artifact hashes and output profile.
What to measure
Metric
Definition to agree
Why it helps
Unresolved content defects
Confirmed omissions, substitutions or repetitions per accepted audio hour
Tracks fidelity more directly than general “naturalness.”
Navigation accuracy
Correct labels and destinations divided by evaluated items, by type
Prevents chapter success from masking note/page failures.
Review effort
Active human review minutes per accepted audio hour
Measures the practical cost of automation.
Repair burden
Attempts and regenerated audio per accepted hour
Exposes unstable models and costly chunk sizes.
Throughput
Accepted output hours per elapsed production hour and per compute hour
Separates machine speed from accepted-book delivery.
Operational reliability
Failed/resumed jobs, incomplete packages and recovery time
Tests production behavior rather than isolated demos.
These are proposed definitions, not measured results or contractual targets. Report the test corpus, input class, language, hardware, model version and accepted-output denominator with every result. Avoid hiding rejected books from quality statistics; report rejections separately.
Playback tests beyond the opening chapter
Exercise all navigation levels, notes/sidebars and skipping behavior, segment stepping, resume, transitions between audio files, and faster playback. Check both protected and unprotected artifacts where relevant. Record the player/app version and profile.
The draft expressly calls for testing with the Library playback environment. The exact device matrix and pass criteria remain open decisions. Accessibility of the operator interface should also be tested with keyboard and assistive technology as a proposed implementation practice.
Keep acceptance independent
Use known reference books, deliberate malformed examples, and a held-out evaluation set. An AI coding assistant can help generate tests and analyze regressions, but the acceptance criteria and reference outputs should be reviewed independently from the implementation that produces them.
The draft requires integration with the Library's AWS systems and standard import/export without proprietary lock-in. The components below are a proposed implementation, not a mandated stack. Confirm available services, network boundaries, security rules and delivery interfaces with the Library. SOW §3.7.
Prototype hardware and cloud boundary
See Mac hardware and cloud options for proposed development tiers, single-GPU AWS workers and an adjustable capacity worksheet. The supplied draft supports AWS integration; it does not provide an allowlist for other providers or settle account, region or service approval.
Proposed components in the Library’s approved AWS environment
Application + control planeBook versions, jobs, approvals and output profiles
↔state
PostgreSQLStores state and decisions for the control plane
↓ schedules jobs and records results
Workflow orchestrationStep Functions + AWS Batch, or an approved equivalent
↓ dispatches bounded, resumable work
Model workersTTS · transcription · alignment Pinned models and voices; GPU or suitable CPU
Compilation workersParsing · XML · media tools · validation · packaging Isolated CPU jobs
↕ workers read inputs and write versioned artifacts
S3 artifact storage
Originals→Intermediates→Staging→Accepted outputs
Delivery adapter Submits only the accepted package; records acknowledgement
↓ crosses the agreed delivery interface
NLS receiving workflowPublic 1206 guidance names the Transfer Portal. The project-specific integration still needs confirmation.
Proposed deployment, not a prescribed AWS stack. The control plane manages decisions, workers process content, S3 stores artifacts, and a separate delivery adapter submits accepted books. S3 storage is not itself the confirmed submission interface. The diagram shows logical responsibilities, not a network-security design.
Proposed deployment components
Component
Candidate implementation
Responsibility
Control plane
Typed application service
Book versions, jobs, profiles, approvals and audit events.
Metadata and job state
PostgreSQL
Structured records, transactions and searchable status.
Artifact storage
S3
Separate originals, intermediates, staging and accepted outputs.
Speech and analysis workers
Containerized Python workers with pinned models
GPU inference or appropriate CPU inference.
Compilation workers
Isolated CPU containers
Parsing, XML, media inspection and packaging.
Workflow orchestration
Step Functions with AWS Batch, or approved equivalent
Dependencies, bounded retries, checkpoints and review gates.
Operator application
Accessible web interface
Exception review, progress and authorized release.
Delivery adapter
Agreed NLS integration
Submit validated packages and record acknowledgement.
Use immutable inputs, explicit output versions and idempotent job keys. A retried job should not duplicate a submission or overwrite an accepted book. Record checkpoints after durable artifacts are written and verified. Use separate states for failed processing, unresolved content, ready for review, and ready for delivery.
Keep model workers warm over a batch where appropriate. Loading weights for every paragraph wastes time, but the right worker lifetime depends on measured throughput and utilization.
Capacity without invented benchmarks
A useful planning equation is:
GPU-hours of work = source audio-hours × processing factor × retry multiplier
Wall-clock hours ≈ GPU-hours of work ÷ (GPU count × effective utilization)
Here the processing factor must be measured in GPU-hours per source audio-hour for the chosen workload. If it is a combined factor, avoid counting transcription or repair twice.
For illustration only: 30,000 hours × 0.5 GPU-hours per hour = 15,000 GPU-hours before retries. At a 1.2 retry multiplier, this becomes 18,000. Ten GPUs at 75% effective utilization would imply about 2,400 elapsed hours of compute processing. These are assumptions and arithmetic, not observed performance or a calendar delivery forecast. Human review, CPU work, queues and delivery can remain bottlenecks.
Benchmark the two input paths separately. Commercial recordings may need full-book transcription/alignment but only supplemental TTS; electronic text needs full narration. Measure accepted output, including failed attempts, rather than quoting isolated generation speed.
Security and source handling
The draft references federal controls in Section H, which is absent from the supplied attachments. It also calls for source/output protection, retention/destruction procedures and security documentation. Obtain that missing material before fixing the controls baseline. SOW §§3.6, 3.13.1, 5.1.
Proposed defaults include isolated inference in the approved environment, restricted unnecessary egress, pinned dependencies, least-privilege roles, separate content-protection keys, and limited source text in logs. Book content must be treated as data, including text that looks like instructions. SOW §5.1 limits use of Government source materials to contract performance and requires secure disposal at contract end. Implement those obligations alongside the agreed retention and deletion procedures.
Storage encryption, transfer encryption and PDTB protection solve different problems. Keep their responsibilities and key handling explicit.
Handover and maintenance
Deliver source code, deployment instructions, configuration, dependency/model/voice inventory, operator and administrator guides, recovery procedures, validation tools and a known reference build. Confirm which third-party components can be redistributed and which must be supplied by NLS.
The draft specifies support response within 24 hours and a 12-month defect-correction warranty after acceptance; it permits ongoing support/maintenance separately from prohibited recurring access requirements. See the milestone page and SOW §§3.8–3.9, 3.12.
Release evidence
For each released software/profile version, preserve migration notes, a reproducible deployment recipe, dependency hashes, regression results and the reference-book outputs. That makes later maintenance possible without relying on the original contractor's environment.
Use technical proof gates to reduce uncertainty while meeting the draft's separate contractual phases. The five gates below come from the engineering proposal; they are not the four phases named in the SOW.
Detailed prototype work plan
The prototype implementation plan adds a proposed 16-week engineering sequence, scope boundaries, owner roles and handover artifacts. The acceptance plan defines the proposed corpus and evidence. These planning estimates do not replace the draft deadlines below.
Proposed proof gates
Gate
Work to demonstrate
Pass evidence
1. Format and protection
Build a tiny complete book with the approved profile, encoder, protection and packaging
Approved validation results and working navigation/playback on target systems.
2. Long-form speech
Compare candidate engines on chapters and complete representative books
Agreed fidelity and quality results, with measured correction effort.
3. Both ingestion paths
Process structured text, less-structured text, commercial audio with and without matching text
Source coverage and independently checked navigation, including explicit unavailable information.
4. Repair and rebuilds
Change pronunciation, replace a paragraph, edit a heading, interrupt processing and fail encoding
Only necessary work repeats; all dependent timing, hashes, metadata and packaging remain correct.
5. Production operation
Run batches, security checks, recovery and real integration
Repeatable accepted outputs, measured throughput and operational handover.
Prove codec/protection compatibility early with simple known content so that format errors can be separated from speech errors. Speech and importer experiments can proceed in parallel after their required inputs are available.
Schedule in the supplied draft
POP means period of performance. No actual start date or total duration is supplied here, so the dates below are relative triggers rather than calendar commitments. SOW §4.1 and named phase descriptions.
Deliverable or event
Draft timing
Key personnel resumes
At proposal submission.
Kickoff
Within 10 business days; the table uses POP start while §5.2 uses contract award. Clarify the trigger.
Five architecture diagrams and eight technical documents
Within 30 calendar days after POP start.
Status and quality-control reports
Every 30 calendar days during POP.
Phase I prototype, sample DTBs, workflow demonstration and test report
Within six months of POP start.
Phase II integration and operational testing
Starts after Phase I completion and Government acceptance; completed within 90 calendar days.
Phase III full deployment
Starts after Phase II completion and Government acceptance; completed within 90 calendar days.
Phase IV ongoing support
After Phase III completion and acceptance, through the remainder of POP.
Warranty
12 months after written Government acceptance of all deliverables.
The 24-hour support-response requirement appears in §3.8. The warranty includes free defect correction and patches under §3.9. Do not turn the phase durations into a promised total project length: Government acceptance timing and the total POP are not established here.
How the gates can fit the phases
Proposed mapping: use Phase I to resolve the profile and prove format/protection, model behavior and initial ingestion. Continue hardening ingestion, repair and operational behavior through Phase II integration tests. Phase III adds deployment, handover and acceptance evidence. Phase IV covers production support and controlled improvement.
This mapping does not reduce the deliverables required in any phase. Validate it against the final solicitation and approved project plan.
Early documents should expose real decisions
The first-month architecture package should show both source paths, data ownership, security boundaries, failure/retry handling, and the split between AI proposals and accepted compiler inputs. List unresolved standards and supplied-component dependencies rather than hiding them behind a generic cloud diagram.
Suggested readiness check before a pilot
Confirm the output profile; acquire reference books, target players and required tooling; agree supported content classes and quantitative acceptance criteria; record authorized voices and dependency rights; obtain Section H; and identify who can accept exceptions and release a book. These are suggested prerequisites tied to the decision register.
The items below remain unresolved by the supplied materials. Suggested decision partners describe the kind of authority needed; they are not assignments made by NLS. Do not mark an item resolved without a written answer, accepted test result or other attributable evidence.
Decisions that affect the architecture
Question
Why it matters
Evidence needed and suggested decision partner
Which editions and precedence govern the 2002/2005 conflict and current-version language?
Changes XML structures, limits and test rules.
Approved profile matrix from the contracting/technical authority.
Which multimedia types are required at acceptance?
audioNCX and synchronized full text imply different outputs.
Supported output-type list and reference books from NLS production/QA.
Which encoder, protector, validators and keys will NLS provide?
Determines the feasible encoding/protection route and deployment platform.
Current packages, versions, terms, headless interfaces, OS support and cloud approval.
What accompanies commercial recordings?
Metadata and matching text determine recoverable structure and pages.
Representative source sets and expected navigation from NLS production.
How do mastering rules apply to synthetic audio and generated labels?
Native model rate, initial-recording language and copied heading rules may conflict with a chosen workflow.
Approved interpretation and accepted synthetic example from NLS audio/QA.
What security controls and retention rules apply?
Determines network, access, key, logging and deletion design.
Missing Section H and approved architecture requirements from Library security.
What constitutes an accepted book?
Controls quality thresholds, human review and pricing/effort assumptions.
Content classes, language scope, player matrix, defect thresholds and throughput/turnaround targets.
What is the final submission interface for this project?
The public Transfer Portal and proposed internal S3 staging are different interfaces.
Integration contract, credentials/process ownership and acknowledgement semantics.
How are third-party model, voice and component rights handled?
Library ownership/non-subscription terms must coexist with component rights.
Approved inventory and applicable license/voice permissions from appropriate procurement/legal reviewers.
Specific mastering clarification
The live public mastering page is 1202:2026, revision 1.5, June 3, 2026. It specifies initial 44.1 kHz PCM and adds integrated loudness of −22 LUFS ±2; it also says heading clips are copied from primary audio. The public index still advertises a 2025 entry. Confirm the actual governing edition and how these rules apply to neural synthesis and separately generated navigation speech. Resampling a model's output alone does not settle the “initially recorded” interpretation. Current mastering page and specification index.
Additional source gaps
The SOW mentions 1201; J1 references 1208 for commercial-audio masters; only 1203 is among the four attachments. Request the complete applicable specification set. The phrase “and D standards” in SOW §3 is unexplained. The kickoff start trigger differs between §4.1 and §5.2. These are useful clarification items, not grounds to invent a preferred interpretation.
Proposed decision record
For each answer, store the exact question, governing source and section, answer date, responsible authority, evidence link, affected profile/version, affected tests and remaining conditions. Preserve superseded decisions as history.
Example: approving a new file-count rule updates the output profile and the corresponding validator. It should not quietly change an already submitted book or erase the rule used for its original build.
What can proceed while answers are pending
Build the source-preserving data model, importer experiments, independent quality checks and model evaluation harness. Prototype against a clearly labeled provisional profile. Avoid committing to a production codec/protection stack or promising numerical quality/throughput until the relevant components and acceptance rules are confirmed.
This glossary explains the terms as used in this proposed production system. The linked articles and standards contain the implementation detail.
Term
Meaning
NLS
The Library of Congress National Library Service for the Blind and Print Disabled.
DTB
Digital talking book: audio and supporting structure that allow accessible reading and navigation.
DAISY
Standards and technology for accessible publications, including structured talking books.
ANSI/NISO Z39.86
The talking-book standard cited in the procurement; the edition matters.
audioNCX
A multimedia type describing full audio with navigation structure; it does not by itself promise a synchronized full-text product.
TTS
Text to speech: generating spoken audio from text.
ASR
Automatic speech recognition: generating a transcript from audio.
Forced alignment
Estimating where a supplied transcript fits the audio; distinct from independently recognizing what was spoken.
Canonical book model
The shared structured record linking original sources, approved text, navigation, audio and evidence.
Provenance
Records of where an item came from and how it was transformed.
Source span
An identifiable passage or audio range in an original source.
Spoken form
The approved text representation intended for narration, such as expanding a currency amount.
Lexicon
A book-level or shared set of pronunciation decisions.
Inference
Running a trained model to generate or analyze content.
Model weights
Numerical parameters of a trained model; they are distinct from application code and voice assets.
Deterministic compiler
Conventional software that produces prescribed files from explicit accepted inputs and a versioned rule set.
Output profile
The exact standards, media, metadata and packaging rules chosen for a build.
OPF
The package document describing metadata, included resources and organization.
NCX
Navigation Control file for XML; defines navigation structures and labels.
SMIL
Synchronized Multimedia Integration Language; links playback structures to media segments.
DTD
Document Type Definition; a schema mechanism used by these XML formats.
Entity
An XML declaration/reference whose allowed definitions must be controlled and packaged as required.
navPoint
A hierarchical navigation point, such as a chapter.
navTarget
A target in an auxiliary navigation list; its allowed use depends on the standard edition.
Segment
A useful navigable unit that need not be a chapter or heading.
PCM / WAV
PCM represents uncompressed audio samples; WAV is a common container for those samples.
Sample position
An integer location in audio at a known sample rate, useful for precise internal timing.
AMR-WB+
The extended wideband audio codec required by the cited NLS profile; different from AMR-WB.
3GP
The container used for the compressed audio, with NLS-specific constraints.
LUFS
A unit for measured perceived loudness; the applicable mastering edition defines the target.
PDTB
Protected Digital Talking Book; content protection and authorization for the book/player ecosystem.
MD5
A checksum algorithm used in the specified DTB checksum format. A checksum is not an authorization or authenticity system.
CRC
A ZIP integrity check value; the delivery specification constrains how it appears in the headers.
Idempotent job
A job whose retry does not create unwanted duplicate effects.
Incremental build
Regenerate changed outputs and their dependencies while reusing unaffected accepted artifacts.
S3
AWS object storage; proposed for internal source and artifact storage here.
GPU
A processor commonly used to accelerate model inference.
Real-time factor
Processing time divided by audio duration under defined hardware and workload conditions.
POP
Period of performance: the contractual time window for performing the work.
SOW
Statement of work: the document defining requested work and deliverables.
RFI
Request for information: a market-research request, distinct from an award or accepted implementation.
For precise format definitions, use J2 or the 2002 standard according to the approved profile. General operational definitions in this table are explanatory additions.
Sources were checked for this local edition on October 1, 2026. Attachment citations link directly to the original SAM.gov downloads. Preserved local copies remain available as backups; live links may change. Nearby article citations identify the source of specific claims. Public references do not automatically supersede attachments or resolve procurement ambiguities.
Original opportunity and attachment context. Check it for later amendments.
The attachment links above are public SAM.gov download endpoints, without session tokens or API credentials. If a direct download becomes unavailable, return to the opportunity attachment list or use a local backup.
The original analysis and SOW text extract are local supplementary files, not attachments hosted on SAM.gov. The SOW text extract is a convenience copy; layout and table relationships may be flattened, so use the Word original to resolve ambiguities.
PDF links with #page= use one-based file page numbers, not necessarily the printed number on the page. Browser support for opening at a specific PDF page varies, and SAM.gov may download a PDF rather than display it at the requested page.
Catalogued November 2007 reference; direct retrieval failed during this review. Its clauses and governing revision were not independently verified here.
AWS Batch GPU jobs and Step Functions with Batch document orchestration options. The Library's approved environment and actual workloads determine whether those options fit.
What changed from the original analysis
The wiki reorganizes the analysis into linked topics and adds explanation rather than treating every recommendation as a requirement. Source review added resource-audio timing exceptions, WAV-relative timing semantics, exact attached-edition limits, draft milestones, support/warranty terms, a missing-specifications list, the live mastering edition, and the distinction between S3 staging and the public Transfer Portal.
The original analysis remains available unchanged for comparison. No benchmark, cost quote, accepted architecture, contract interpretation or playback demonstration has been added as a claimed result.
Prototype planning references
The prototype adds an engineering plan, hardware allocations and test targets, all labeled as proposals. Primary sources below document model/runtime capabilities and available hardware, not NLS approval or measured performance. Checked October 1, 2026.