Talking Book Field GuideNLS · Engineering wiki
Start here

Digital talking book production wiki

Understand the proposal, choose a reading path, and find the evidence.

Download this article as MarkdownArticle permalink

This wiki explains a proposed system for the Library of Congress National Library Service for the Blind and Print Disabled (NLS). It would turn electronic books and existing commercial audiobooks into accessible digital talking books (DTBs) with reliable navigation, required audio formats, protection, and delivery packaging.

The central recommendation is an AI-assisted production pipeline with a deterministic compiler at its center. AI proposes structure, prepares speech, and helps detect defects. Conventional software enforces the approved rules and creates the final files. The original analysis considers this feasible; no model benchmark, compliant prototype, or successful NLS-player demonstration is established by the materials here.

Build the prototype

Start with the prototype implementation plan, then compare the local model stack, Mac/AWS hardware and sizing worksheet, and test corpus and acceptance gates. This is a proposed plan; no model benchmark, hardware purchase or prototype implementation is recorded here.

Explore the diagrams

Choose a reading path

Your goal Read these pages in order
Understand the opportunity Requirements → Architecture → Open decisions
Evaluate the technical approach Book model → Ingestion → Speech models → Compiler
Plan a pilot Prototype plan → Models → Hardware → Acceptance tests
Estimate operational effort Quality review → AWS and operations → Open decisions
Check a claim Follow its nearby source link or visit the source library.

Follow the production path

Two input paths feed one common book model. The clickable flow below opens the corresponding explanations.

How to interpret the claims

Source requirement means a statement found in the supplied draft SOW or a named specification. Edition and applicability still matter. Proposal means an engineering choice from the analysis or this wiki. Open decision means an answer is needed from NLS or from testing before the design can be fixed.

Public specifications are useful comparison material. A newer public edition does not by itself amend this opportunity. The standards page explains the 2002/2005 mismatch and the difference between the attached and current public construction specifications.

What to prove first

Produce a tiny, complete, protected book that passes the approved validators and plays correctly on the intended NLS systems. In parallel, test sustained narration on representative complete books and measure correction effort. These demonstrations address format interoperability and speech quality separately.

About this local edition

Prepared on October 1, 2026, from the supplied analysis, the four opportunity attachments, and linked primary sources. Added explanations include a glossary, explicit acceptance evidence, dependency tracking, milestones, and decision prompts. Nothing in this wiki records answers from NLS or actual performance results.

Wiki content works offline. Attachment citations open the original files on SAM.gov and require an internet connection. The source library also offers local backup copies. Editable articles are stored alongside it as Markdown. Keep the whole folder together when moving it. See the source library for provenance and the glossary for unfamiliar terms.

Build the prototype

Prototype implementation plan

A scoped build, proposed sequence, staffing, deliverables and decision gates.

Download this article as MarkdownArticle permalink

Proposed baseline: build a working, inspectable system that converts electronic text and supplied commercial audio into sample NLS digital talking books, with an operator able to correct a defect and rebuild the affected output. Start on an Apple Silicon Mac; prove the same workflow on one GPU in an approved AWS development environment. The proposed narration engine is Qwen3-TTS 1.7B CustomVoice, with Kokoro-82M as a compact comparison baseline; model selection remains conditional on listening tests and rights review.

Status: plan only. The wiki and its hosting are implemented; the production prototype, model benchmarks, purchased hardware, and NLS acceptance are not. Scope, staffing, durations, memory budgets and pass thresholds below are engineering proposals. Contract statements refer to the supplied draft SOW, not an awarded or complete contract.

Read model choices, Mac and AWS hardware, and test corpus and acceptance alongside this plan.

What the prototype must prove

The prototype should answer four questions: can the chosen voice read complete books faithfully; can the compiler produce a protected book that works on the intended players; can operators resolve exceptions without redoing a whole book; and can the Library operate the delivered system without a required paid speech API or recurring software-access license?

Phase I in the draft calls for an initial working TTS prototype, DTB samples, voice samples, an end-to-end workflow demonstration and a test report. The broader SOW covers both input routes. This proposed prototype deliberately exercises both so commercial-audio risks appear before Phase II. It does not imply that Phase I alone completes all integration, production or support obligations. Draft SOW §§3.1–3.2, 3.12–3.13; deliverable table §4.1.

Scope and boundaries

Area Included in this prototype Boundary or follow-on work
Text input EPUB, DOCX, RTF and UTF-8 plain text; preserve reading order, headings, source spans and metadata Confirm whether “Word Documents” includes legacy .doc; quarantine unsupported, malformed, encrypted or DRM-protected inputs rather than silently drop content. Scanned-page OCR is a separate extension.
Commercial audio WAV, MP3, OPUS and FLAC fixture coverage; preserve narration, propose navigation from available metadata/text/ASR, synthesize approved announcements Do not recreate the narrator or invent unavailable page references. Audio-only boundaries need review.
Languages and content English prose first, two approved selectable voices; ordinary headings, lists, notes, numbers and abbreviations English is a proposed pilot limit, not an SOW exemption. Mixed-language text, equations, poetry, complex tables and illustrations must be flagged and given an agreed reading treatment. Wider language support is a follow-on gate.
Output One agreed audio-plus-navigation profile, WAV masters, required XML/metadata, AMR-WB+, PDTB protection and delivery package MP3 is optional in the draft. Other multimedia types remain a scoped decision. A WAV/MP3 demo or unprotected package cannot substitute for required protected-output proof.
User workflow Import, inspect structure, select voice, queue job, inspect issues, edit pronunciations/labels, regenerate, compare versions, approve and export One production team and simple roles; no multi-tenant service, billing or public upload portal.
Automation Three-book batch, durable checkpoints, bounded retries, cancellation, resumable work and dependency-based rebuilds No promise to sustain the approximately 3,000-book annual background volume until representative benchmarks exist.
Deployment Native Mac model worker, shared API/data contracts, Linux NVIDIA worker in approved AWS test environment Real Library storage/distribution integration belongs to the agreed interface plan and Phase II acceptance. Prototype adapters may be simulated, visibly labeled.
Handover Source, pinned models and voices where redistribution is permitted, deployment recipe, test evidence and operator/admin instructions Production HA/DR, large-scale load qualification, full training and support remain later-phase work; backup/restore and access controls still receive prototype tests.

Build sequence and stop conditions

  1. 1
    Fix the test boundaryAgree profile, lawful fixtures, voices, tools and player access.
    ↓ supplies rules and reference evidence
  2. 2
    Make one tiny protected book playProve XML, timing, encoder, protector and real navigation together.
    ↓ establishes a working output path
  3. 3
    Complete both input routesFull narration from text; preserved commercial narration plus approved speech.
    ↓ provides complete books to inspect
  4. 4
    Measure, repair and recoverListen, flag defects, rebuild dependencies and survive interrupted jobs.
    ↓ produces qualified recipes and evidence
  5. 5
    Repeat in approved AWS and hand overDemonstrate reproducible operation; submit evidence for authorized acceptance.
Each gate earns the next claim. Import and model experiments may run in parallel; missing protection tooling or player access leaves the corresponding output gate blocked.

Planning allowance: 16 engineering weeks, assuming roughly four full-time engineering contributors plus fractional project management/specialists, available source rights, an NLS technical contact and timely access to required tooling and players. This is an internal estimate, not a delivery promise. Calendar commitments must follow the actual POP start and Government decisions. The draft sets a six-calendar-month Phase I deadline; do not equate that deadline with 24 weeks or assume acceptance/review latency is zero. Draft schedule.

Workstream and proposed window Concrete work and owner role Exit evidence / dependency
Weeks 1–2: baseline and risk probes Technical lead + audio/ML engineer: requirements matrix, provisional profile, corpus rights, two-voice short sample, protected tiny-book spike, Mac runtime probe Recorded decisions; encoder/protector/player access requested immediately. A missing codec/protection route blocks a claim of complete DTB success.
Weeks 1–4: design package Technical lead + security/QA contributors: five specified diagrams and eight specified documents Submit within 30 calendar days of POP start, regardless of the week labels. Unresolved decisions remain explicit.
Weeks 3–6: common core Backend/DTB engineer: source store, versioned book schema, EPUB/DOCX/RTF/text import, compiler rules and job state machine Tiny reference book round trip; no unaccounted source text or broken navigation; stable worker contracts.
Weeks 5–9: narration and review ML engineer + product/backend engineer: model adapters, pronunciation overrides, issue queue, transcript/audio comparison, versioned approvals Full text-origin book, both voices, deliberate pronunciation repair with correct downstream invalidation.
Weeks 7–11: commercial path Backend/DTB + ML engineers: media decoding, transcript alignment, metadata navigation, supplemental speech Complete commercial-source sample with matching text; audio-only sample with uncertainty surfaced for review.
Weeks 10–13: reliability and evaluation QA/accessibility lead + all engineers: complete corpus, malformed inputs, crash recovery, bounded retries, independent playback Measured defect and repair rates, all output checks, corpus coverage, accessible operator flow.
Weeks 13–16: AWS proof and handover Technical lead + platform/security contributor: approved AWS sandbox, controlled model transfer, integration rehearsal, clean rebuild and demo Same reference suite passes; compute/storage/review costs recorded; handover rehearsal and test report.
Remaining time before Phase I deadline Project lead coordinates Government feedback, component delays and corrective work Preserve contingency explicitly. If required access is late, revise the schedule and report the blocked gate rather than relabel a partial demo as accepted.

Some work overlaps; format/protection proof, approved output rules and lawful voice/model use constrain later release claims. Model experiments and source-preserving import work can continue while external answers are pending.

Concrete implementation shape

Use a modular application before distributing every component. Proposal: Python API and workers, PostgreSQL for state, an accessible browser operator interface, filesystem artifacts on Mac, and S3 artifacts in AWS. A small job queue and worker leases are sufficient initially; production orchestration can use approved AWS Batch/Step Functions after the basic contracts are stable. See architecture and operations.

Module Prototype behavior Durable contract / output
Importer Validate file type, sandbox document/media parsers, preserve originals, extract structure and source spans Source hash, parser version, ordered nodes, warnings, catalog claims with provenance
Book service Version source interpretation, spoken-text decisions, pronunciation rules, voice and output profile Immutable book version; every accepted passage mapped to a source span or an approved announcement
Job service Enqueue stage jobs, lease work, checkpoint, cancel and resume Input hash + recipe + stage identity; explicit failed/review/released states; no duplicate delivery
Speech worker Generate one bounded chunk under a book-level voice recipe PCM audio, duration, model/voice hash, runtime/precision, parameters, peak memory and elapsed time
Analysis worker ASR, alignment, silence/clipping checks and text-difference proposals Issues tied to exact text/audio versions; no autonomous acceptance
Compiler Assemble accepted masters; derive timing; build XML, encode, protect and package Profile version, actual delivered files, validator report, checksums and dependency graph
Review and delivery Human resolves exceptions; release requires current evidence; delivery adapter records receipt Actor, decision, rationale, approved version, package identifier and acknowledgement

Run local Metal/MLX inference as a native macOS process. Keep database/API containers optional. The Linux CUDA image is a separate build behind the same worker interface, not a claim that a Mac container image has GPU parity. Platform-dependent numerical differences are expected; compare coverage, quality and output validity rather than require identical generated waveforms.

Operator demonstration

Demonstrate a traceable session: import an EPUB, inspect recovered chapters and metadata, choose an approved voice, generate a chapter, inspect a flagged word, correct its pronunciation, audition before/after, rebuild the affected audio/timing/package, run validators, and navigate the protected result on a target player. Then repeat with a commercial audiobook, preserving its narration and adding only the approved generated material.

The UI needs book/job lists, a navigation-tree editor, text/audio issue review with keyboard controls, voice and lexicon settings, an evidence panel and a release action. A release blocked by stale validation should explain which change invalidated the evidence. Separate operator, reviewer/releaser and administrator permissions even if one person holds several roles during the pilot.

Deliverables and staffing

Propose four core contributors: a technical lead/platform engineer, a backend/DTB engineer, a speech/ML engineer and a QA/accessibility engineer. Product/UI design, audio editorial expertise, Library integration/security and procurement/license review are fractional support roles; they are real dependencies, not assumed free capacity. Reserve NLS reviewer and player access in the schedule. With one or two engineers, revise scope or time rather than inherit this estimate.

The day-30 design package should map directly to the SOW's five diagrams: System Content Diagram (wording in the draft), Component Architecture, Data Flow, DTB Packaging/Output Structure, and Security/Access. Its eight documents cover System Architecture, DTB Integration, Text Processing, Audio Output, Metadata/Packaging, Security/Data Protection, Testing/Validation, and Implementation Schedule. The wiki figures explain the proposal; they do not by themselves constitute all these signed-off deliverables. SOW §3.13.1 and §4.1.

Map the technical lead to the Senior Engineer key-personnel responsibility and allocate a separate Project Manager, proposed at 0.25–0.5 FTE during the prototype. The draft identifies both roles in §6.1; actual personnel must satisfy the stated qualifications. The PM owns the decision log, dependency escalation, schedule and status/quality-control reports every 30 calendar days. This management capacity is additional to the four engineering contributors, not hidden in their estimates. SOW §§3.10–3.11, 4.1, 6.1.

End-of-prototype delivery includes a reproducible code release, dependency/model/voice inventory and license notices, configuration/profile exports, corpus manifest, voice comparison, sample DTBs, test results, benchmark raw data, known limitations, backlog, setup/rollback/restore runbooks, a workflow recording and an operator/admin walkthrough. No source books, private recordings or keys go into the public wiki.

Decisions before implementation spending

Resolve the output-profile editions, synthetic-audio mastering interpretation, available encoder/protector and OS support, authorized voices, Library account/region/network access, Section H controls and retention, target players and acceptance owner. Obtain throughput and human-review budgets before selecting production capacity. Decision register.

If an encoder or protection tool is Windows-only, add a controlled Windows worker and a license/deployment review; Mac development does not remove that dependency. If full-book speech quality fails, compare the baseline and challenger or narrow the agreed content class; do not hide failures behind unbounded retries. If a model license cannot meet the ownership/access terms, replace the component before investing in its integration.

Build the prototype

Proposed local model stack

Qwen narration, comparison models and practical Mac-to-CUDA runtime choices.

Download this article as MarkdownArticle permalink

Start with Qwen3-TTS-12Hz-1.7B-CustomVoice for narration, and evaluate it against Kokoro-82M. This is a proposed starting point, not a finding that either model meets NLS requirements. Keep model inference local to the developer Mac or the approved AWS environment; the delivered workflow should not depend on a paid speech API. Training or fine-tuning a new model is outside the initial prototype scope.

Primary narration choice

The exact upstream artifact is Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice. Its card lists Apache 2.0, preset voices, multilingual synthesis and style instructions. Choose CustomVoice, rather than Base, for the first prototype because it provides preset speakers without requiring a new voice-cloning workflow. The official examples use CUDA/BF16. Official Qwen implementation.

Start with restrained narration and short sentence-group chunks. Upstream lists Ryan and Aiden as English-native presets, both male. Include other approved English voices in the comparison if the pilot needs a broader selection; multilingual capability alone does not prove equal narration quality for every speaker/language pairing. Treat speaker choice as book-level configuration. Never silently switch engines or voices to repair one sentence.

The official checkpoint tree is approximately 4.52 GB at review time, including associated files. That is download size, not peak inference memory. The “1.7B” name alone cannot determine the total working footprint. Pin the selected revision and record actual memory during the runtime probe.

Supporting models and alternatives

Role Proposed artifact What it does, and its limit
Compact TTS comparison hexgrad/Kokoro-82M Small Apache-2.0 model with 24 kHz sample output. Useful for plumbing, announcements and voice comparisons; English voice catalog provides more preset choices. Its size does not establish full-book quality.
Independent speech check openai/whisper-large-v3-turbo MIT, 809M-parameter ASR candidate. Transcribes generated and commercial audio to flag discrepancies. Hallucinations and recognition errors mean its transcript is evidence for review, not truth.
Alignment experiment Qwen/Qwen3-ForcedAligner-0.6B Apache 2.0; aligns supplied text/audio, with a documented maximum of five minutes per input. Use much shorter overlapping windows and reconcile boundaries. Alignment cannot independently prove the supplied words were spoken.
Optional structure assistant Qwen/Qwen3-4B-Instruct-2507 Apache 2.0 instruction model for schema-constrained heading/reading-order suggestions on ambiguous input. Use a proposed 4–8K context initially. It must not rewrite narration, invent catalog facts or generate authoritative final XML.
Later challenger Chatterbox MIT project with several model variants and an embedded watermark. Compare only if the initial two-engine evaluation leaves a meaningful quality gap; validate exact variant, voice rights, runtime and watermark behavior after encoding.

The prototype needs TTS plus independent audio checking; it does not need every optional model loaded or even integrated on day one. Deterministic parsing remains the first choice for documents with usable structure.

Mac and AWS runtime paths

Stage Apple Silicon development Linux NVIDIA reference
Qwen narration MLX-Audio Qwen adapter, beginning with the documented BF16 CustomVoice conversion Upstream Qwen PyTorch/CUDA runtime; BF16 baseline with pinned dependencies
Kokoro comparison Upstream Kokoro using CPU or MPS as supported by the pinned version Same upstream model and code on CUDA/CPU
Whisper discrepancy checks whisper.cpp, Apple Silicon/Metal, with a pinned model conversion faster-whisper, with matching model source and recorded decode/precision settings
Optional aligner MLX-Audio conversion only after timestamp checks against known fixtures Upstream Qwen aligner as reference implementation
Optional structure LLM MLX-LM, evaluated 4-bit conversion Pinned local inference runtime and the same approved source model; benchmark separately

MLX-Audio is a community implementation. Its support is not a claim that the original CUDA package runs unchanged on Mac. Qualify the conversion on identical passages and on full-book continuity. Metal/MLX uses the Mac's shared memory; a CUDA worker has separate GPU and host memory. Hardware tiers.

For the first quality comparison, keep Qwen TTS at BF16. Quantized TTS is a separate, held-out experiment. Record runtime-specific output changes; do not assume a smaller checkpoint preserves fidelity. Use separate inference environments when dependencies conflict.

Guardrails in the adapter

Each synthesis request carries approved text, language, speaker ID, narration style, pronunciation policy, model and runtime versions, precision and a deterministic job identifier. Save audio duration, sample rate, text hash, elapsed time, memory peak and warnings with the output. Seeded generation improves traceability but does not guarantee identical audio across hardware.

The adapter must reject empty or incomplete output and detect truncation. Kokoro's pipeline implementation has bounded phoneme sequences and truncation-related behavior; enforce complete source coverage before calling inference, not merely by reading logs afterward. Use sentence boundaries and source spans to split material, and keep the distinction between synthesis chunks, navigation units and packaged files.

Pronunciation overrides change a separately recorded spoken representation. Do not feed arbitrary SSML to a model unless its adapter explicitly supports it. Check names, numbers, units and negation even when overall ASR word error rate looks good. Use the test plan for release evidence.

Rights, mastering and handover

Apache/MIT labels are useful starting evidence, not blanket clearance of model provenance, voice rights, dependencies, redistribution or the draft's ownership/non-infringement terms. Maintain a reviewed inventory of exact checkpoint revisions, conversions, voice assets, code licenses and any reference-speaker permission. Voice cloning, automatic character voices and style effects are deferred unless explicitly approved for the content.

Pre-fetch approved artifacts into a controlled cache, verify hashes and prove inference with network access disabled. Continued operation must not require a vendor account, hosted model endpoint or license renewal. The final rights arrangement still needs procurement review against SOW §§2, 3.9, 3.12.

Native synthesis sample rates and token rates are not output-compliance settings. Keep original generated audio, then apply the approved mastering/encoding profile. Upsampling to 44.1 kHz does not settle whether a model satisfies a requirement about initial recording; resolve the mastering question before making a compliance claim.

Build the prototype

Mac hardware and contract-aligned cloud options

Mac tiers, AWS worker sizes, cloud-permission boundaries and a sizing worksheet.

Download this article as MarkdownArticle permalink

Proposed starting configuration: an actively cooled Apple Silicon Mac with 64 GB unified memory and a 1 TB SSD, plus a short-lived AWS g6.2xlarge test worker with one NVIDIA L4, 24 GB GPU memory, 8 vCPUs and 32 GiB host RAM. Run model stages sequentially at first. This is a practical provisioning proposal for the model stack, not a measured minimum, throughput guarantee or purchase order. AWS G6 specifications.

What the draft establishes about cloud providers

The supplied draft says to integrate with NLS AWS systems (§3.7), and describes existing S3-based storage (§1). It does not provide a cloud-provider allowlist or explicitly require every workload to execute exclusively in AWS. Because Section H is absent, the permitted account, region, services, data residency, network boundary and security authorization cannot be determined from these attachments alone. Draft SOW §§1, 3.6–3.7, 5.1.

Environment / provider Evidence of permission Proposed use
Library-approved AWS environment AWS integration is explicit; particular compute service/account/region approval is not established Primary candidate for cloud inference, storage and integration testing. Confirm service availability and GPU quota before booking work.
AWS GovCloud Not named in the supplied draft A conditional AWS option only if the Library specifies that boundary; do not assume it is required, sufficient or interchangeable with commercial AWS.
Azure, Google Cloud, Oracle Cloud No authorization established by the supplied attachments Do not include as approved runtime choices. Request a written determination before any contract-content processing or deployment there.
RunPod, Lambda, CoreWeave or other GPU hosts No authorization established by the supplied attachments Potential technical alternatives only after provider and data-handling approval; excluded from the baseline.
Developer Mac Proposed engineering environment, not an expressly approved place for Government content Public-domain, synthetic or explicitly authorized fixtures until content-handling permission is confirmed.
This Cloudflare Pages wiki Public documentation hosting already requested by the user Documentation only. Its deployment establishes no permission to host source books, protection keys, private recordings or inference there.

Cloud certifications alone would not establish contractual permission. Keep software-access rights and consumption-based infrastructure costs as separate questions; the draft's non-subscription language should be clarified for the final procurement arrangement.

Mac development tiers

Tier Proposed hardware Intended use and tradeoff
Reuse existing hardware Apple Silicon, 16 GB unified memory; adequate free disk API/UI/compiler work, Kokoro and short serialized smoke tests. Useful starting machine; not the buying target for sustained Qwen plus QA development.
Practical sequential-worker floor Apple Silicon Pro/Max class, 32–36 GB unified memory, 1 TB SSD One model loaded at a time, bounded chunks/context, few competing apps. Verify actual peak memory and swap before accepting it for whole-book runs.
Recommended desktop Mac Studio, M4 Max, 16-core CPU / 40-core GPU, 64 GB unified memory, 1 TB SSD Comfortable development headroom for speech experiments and application tooling. Apple lists this memory/storage combination; the recommendation is an engineering judgment, not an Apple TTS benchmark.
Portable alternative MacBook Pro, M5 Pro with 20-core GPU, 64 GB unified memory, 1 TB SSD; M5 Max optional Similar memory budget in a portable, actively cooled form. More GPU cores may help, but throughput must be measured on the selected configuration.
Larger workstation 96–128 GB unified memory Defer unless simultaneous model workers, a larger LLM or training becomes an evidenced need. Extra capacity is not required by this proposed scope.

Apple's published configurations: Mac Studio specifications and M5 Pro/Max MacBook Pro specifications. These establish supported configurations, not current inventory, delivery dates or prices. Confirm availability before purchase. An existing 48 GB machine is also a reasonable evaluation workstation; do not replace it before the memory/throughput probe.

The Mac's memory is shared by GPU, OS and applications; it is not equivalent to the same number of gigabytes of dedicated NVIDIA VRAM. Reserve at least 200 GB initially free for the pilot, plus a backup destination. Use native Metal/MLX workers; validate Linux/CUDA separately. Headphones for careful listening, representative NLS playback hardware and an accessible test browser/screen reader are as important to acceptance as the GPU.

Working-memory budget to validate

Process Proposed starting allocation Scheduling policy
Qwen 1.7B BF16 TTS 8–12 GB Batch 1, bounded sentence groups; measure cache/codec/framework peaks
Kokoro baseline 2–4 GB Load instead of Qwen for a comparison run
Whisper turbo ASR 3–6 GB Inspect accepted/generated chunks after synthesis; no required co-residency
Optional forced aligner 2–4 GB Short windows; reconcile boundaries
Optional 4B LLM at 4-bit, 4–8K context 4–8 GB Separate structural suggestion job; no whole-book context assumption
OS, database, API, browser and build tools Reserve roughly 10–16 GB on Mac Limit container memory and competing apps; watch memory pressure

All allocations above are estimates to test, not vendor minimum specifications. They depend on runtime, precision, input length, batches and conversion. Do not add every row as if all models must reside simultaneously. Conversely, do not use raw checkpoint size as total RAM. Reduce batch/context first, then serialize or unload workers before buying larger hardware. Model evidence and runtime choices.

AWS instances and surrounding resources

Candidate Published hardware Prototype role
g6.2xlarge — first choice to test 1 × L4, 24 GB GPU memory; 8 vCPUs, 32 GiB RAM One sequential inference worker; separate CPU compilation if host memory becomes tight
g6.4xlarge Same single L4 / 24 GB GPU memory; 16 vCPUs, 64 GiB RAM More host-side parsing/media/cache headroom; it does not double VRAM
g5.2xlarge — availability fallback 1 × A10G, 24 GB GPU memory; 8 vCPUs, 32 GiB RAM Benchmark against G6 if approved region/quota or quoted cost makes it preferable
g5.4xlarge Same single A10G / 24 GB GPU memory; 16 vCPUs, 64 GiB RAM Host-memory alternative; validate the exact backend independently

Hardware values come from AWS G6 and AWS G5 product tables; GB and GiB retain AWS's published units. Family support does not establish capacity, regional availability or Library approval. Avoid fractional-GPU instances for the first reproducibility baseline. No multi-GPU/H100 reservation is justified by current evidence.

Proposed supporting resources: one 4-vCPU/16-GiB CPU application/compile node (or equivalent approved service), PostgreSQL, a private artifact bucket, a private model cache, encrypted 250 GB scratch/EBS capacity for the GPU node, bounded log retention, and backup storage. The initial CPU/database services may share a development node; isolate them for later production requirements. Use durable storage for checkpoints; instance-local scratch is replaceable. Start with on-demand test windows and shut workers down when idle; test interruption recovery before considering Spot.

Same book contract, different execution environments

Package identical schema versions, text fixtures, voice recipes and output profiles for both platforms. Compare semantic fidelity, timing validity, memory and accepted throughput. Do not require byte-identical stochastic narration across Metal and CUDA. Test the approved encoder/protector on its actual supported OS; if it is Windows-only, budget a separate approved worker rather than assume Linux or macOS support.

Estimate compute and review effort

Change the assumptions to see the capacity effect. Defaults are examples, not observed model speed or an AWS quote. Each worker is a single-GPU instance.

Busy GPU processing32.80 GPU-h
Allocated instance time, total43.73 h
Compute wall time, approximate43.73 h
Human review effort, separate30.20 person-h
GPU instance subtotal onlyEnter a rate

No data leaves this page. Changing worker count changes estimated compute duration, not total work. Human review, CPU work, queues and service limits may prevent linear speedup.

Show the arithmetic and exclusions
Generated = text hours + commercial hours × supplemental %
Attempted generation = generated × generation multiplier
Busy GPU-hours = attempted generation × TTS factor + (commercial hours + attempted generation) × ASR/alignment factor
Allocated instance-hours = busy GPU-hours ÷ utilization fraction
Compute wall-hours = allocated instance-hours ÷ workers
Review person-hours = (text + commercial + supplemental hours) × review minutes ÷ 60
GPU subtotal = allocated instance-hours × entered instance rate

The generation multiplier approximates total replacement-audio work, not the maximum retry count. This excludes CPU compilation, optional LLM jobs, repeated full-source analysis, storage, network, engineering and review costs. It assumes enough independent jobs for the requested concurrency.

Measure before sizing production. Substitute observed per-stage factors, actual regeneration and review effort, and a current region-specific quote.

The worksheet uses explicit assumptions; its defaults are illustrative, not benchmarks. It models one GPU per instance and one selected GPU stage at a time. Generated hours include full text narration and supplemental commercial-audio speech. The generation multiplier covers additional synthesis and checking of replacement audio. Commercial original audio is checked once; repeated full-source analysis, optional LLM work, CPU encoding, queue delays and storage/transfer need additional budget. Effective utilization converts busy processing hours to allocated instance-hours; aggregate wall time divides by worker count, while total paid capacity remains unchanged in this simplified model.

Human review is separate and may dominate elapsed delivery. Pricing is deliberately entered by the reader because region, operating system, purchasing arrangement and date change it. Use a current AWS Pricing Calculator estimate for the approved region, and add CPU/DB, EBS/S3, backups, logs, network egress, player access, engineering and human-review costs. Set an account budget and worker time limit before any cloud experiment; no cloud instances are provisioned by this plan.

Storage sizing and buying decision

As an arithmetic example, 44.1 kHz, 24-bit mono PCM uses about 0.476 GB per hour: sample rate × bytes/sample × channels × seconds. Fifty hours with four retained PCM versions is about 95 GB before originals, models, intermediates and backups. This example is not an approved mastering profile. Stereo doubles that quantity; compression and retention change actual needs.

Buy or reserve only after a 30-minute representative narration/analysis run, a long-chapter run and the protected-output spike establish fit. Capture peak memory, sustained swap, per-stage RTF, reviewer effort and tool compatibility. Scale worker count only after a three-book batch and fault recovery pass. Approve production capacity against accepted audio-hours and required turnaround, not model parameter count alone.

Build the prototype

Prototype corpus and acceptance plan

Complete-book fixtures, measurable evidence, failure tests and the demonstration.

Download this article as MarkdownArticle permalink

These are proposed engineering gates, not Government-approved acceptance thresholds or completed test results. The prototype should produce evidence for both source routes, full-book speech fidelity, protected playback, repair behavior and operational handover. The draft requires quality checks and functioning Library playback; it does not supply numerical TTS quality or throughput targets. SOW §§3.5, 3.11, 3.13.

Corpus with a reason for each item

Use public-domain, created-for-test or explicitly authorized content on developer machines. Record permissions, source hashes, language, duration/word count, expected navigation and known difficult passages. Government-provided materials require the approved handling environment. Proposed corpus: six complete books, about 30–50 source/output audio-hours in total, plus a short adversarial suite. The range is a collection target, not a statement about available books or a statistical guarantee.

Corpus component Coverage Human reference work
Four text-origin books One each in EPUB, DOCX, RTF and plain text; include prose fiction, nonfiction with notes, numeric/abbreviation-heavy material and weakly structured text Establish authoritative reading order, metadata, headings and approved spoken forms; distribute difficult features across books.
Two commercial-origin books One with matching text/metadata; one audio-only with limited metadata Reference chapter boundaries and announcements; explicitly mark unknown page information.
Format fixtures Short valid WAV, MP3, OPUS, FLAC plus all four text formats; corrupt, encrypted and misleading-extension examples Known success/rejection result for every fixture. Avoid lossy re-encoding merely to enlarge the corpus.
Difficult-passage set At least 100 approved excerpts with names, numbers, dates, abbreviations, notes, repeated words, punctuation, long sentences and language switches Reviewer-approved pronunciation and reading policy; hold out 20 excerpts from tuning.
Tiny protected reference book Three chapters, nested navigation, announcement, a note and one known page reference Exact expected metadata/navigation plus target-player checks; use an NLS-approved reference package where available.

Compare two approved voices and two engines on the same difficult-passage set. Select one recipe before the expensive complete-book run, retain the holdout set, and use a second reviewer to adjudicate disagreements. A six-book corpus demonstrates coverage; it cannot establish a universal error rate for all genres or languages.

Exit gates

Gate Proposed pass condition Saved evidence
Source fidelity Every source span accounted for; every omitted, expanded or normalized spoken passage has an explicit rule or reviewed decision Source-to-spoken-text map, coverage report, normalization diff; ordinary punctuation transformations recorded by rule
Book structure All required metadata and expected navigation targets are correct in the fixtures; no broken references or unsupported invented page numbers XML/schema and profile validation, expected-vs-actual navigation comparison
Speech quality Full human listening pass of final encoded/decoded release audio for the six-book corpus, including supplemental announcements; zero known unresolved omissions, hallucinated passages, repetitions or material meaning/pronunciation defects in release candidates Timestamped defect ledger, repairs and reviewer sign-off; report residual uncertainty and minor defects separately
Voice suitability Blind comparison of clarity, stability, fatigue and joins; NLS reviewers agree voice acceptance threshold before scoring Per-voice distributions, sample counts, reviewer disagreement and approval; do not claim a generic MOS score proves acceptance
Timing and mastering Measured outputs satisfy the approved profile's clip boundaries, levels, sample rate and encoding rules Automated measurements plus spot-checked waveform boundaries; timing explanation
Encoding and protection Actual AMR-WB+ plus PDTB/NLS protection using approved tools; output decodes/plays through the authorized path Tool versions, protected package hashes, validator results, authorized-key playback; never include keys in reports
Playback Every navigation entry in the tiny book and all top-level/representative nested entries in long books work on the agreed NLS hardware/application matrix Device/application versions, exact actions and observed results; browser audio preview alone is insufficient
Repair correctness Pronunciation, paragraph-length and heading-label changes rebuild every dependent artifact and preserve unaffected accepted chunks Dependency comparison, new timeline/checksums, old validation invalidated, post-rebuild playback
Reliability Three-book queue survives a worker kill, duplicate request, interrupted upload and restart without lost decisions or duplicate release Job-event trace, recovered artifacts, deliberate fault results; proposed retry cap: two regeneration attempts after the first
Operator accessibility Complete import/review/repair/release workflow by keyboard and agreed screen-reader combinations; focus, labels, issue announcements and audio controls usable Manual task test, assistive-technology versions and findings; formal standards baseline agreed with the Library
Security and handover Unauthorized release/content access denied; logs omit source content/secrets; unapproved egress denied; agreed retention/deletion policy verified across originals, intermediates, caches and backups; backup restores a reviewed book; clean installation reproduces the reference workflow Access, retention/deletion and restore tests, dependency inventory, offline model-loading test and another operator's setup log

“No known unresolved defect” means the agreed review found and resolved defects; it does not prove the model can never produce an error. Full listening is deliberately expensive for this small pilot so the team can measure what an efficient production review policy would miss. ASR agreement alone must never grant release.

Benchmark procedure

  1. Freeze model, voice, text normalization, chunking, precision, runtime, driver/OS and machine configuration. Record exact model revision hashes; do not compare mutable latest labels.
  2. Run cold-load and warm inference separately. Use fixed short passages first, then one complete chapter, then whole books. Record output duration and errors along with elapsed time.
  3. Measure TTS, ASR/alignment, CPU mastering/encoding/protection, queue wait, transfer and human review separately. Report p50/p95 chunk latency, peak memory, swap, failure and repair rates; include sample counts.
  4. Repeat representative runs, including long inputs and interrupted jobs. Count failed generations, discarded audio and repair attempts in cost, rather than measuring only accepted best runs.
  5. Repeat the same corpus on Mac and the chosen AWS instance. A faster machine is useful only if accepted output and repair effort remain acceptable. Qualify quantized and converted model artifacts separately.

RTF (real-time factor) = processing seconds ÷ resulting audio seconds for a stated stage. RTF 0.5 means that stage takes half as long as its audio; RTF 2 takes twice as long. Do not use time-to-first-audio or streaming latency as whole-book throughput. The hardware worksheet is a planning tool until measured inputs replace its examples.

Report accepted audio-hours per paid instance-hour, cost per accepted audio-hour, and reviewer minutes per accepted audio-hour. Suggested feasibility target for discussion: combined automated processing at or below the audio duration on the selected cloud worker, with no sustained memory pressure. This is a proposed test target, not a measured result or an NLS requirement; failure should trigger a sizing/quality tradeoff, not suppression of data.

Traceability and report format

Each requirement row should contain an ID, exact SOW/specification edition and section, implementation version, test fixture, expected result, observed result, evidence URI, defect ID and acceptance authority. Status values should distinguish not tested, passed, failed, blocked and not applicable with rationale. A missing player or protector is blocked, not passed.

The final report includes corpus coverage, raw measurements, voice comparisons, issue severity definitions, resolved and open defects, security/environment exceptions, playback evidence and the next-phase recommendation. Record both successes and excluded material. Acceptance should be attributed to the authorized reviewer; the engineering team's internal gates do not replace Government acceptance.

Demonstration agenda

Allow about 60–90 minutes as a planning assumption: inspect the corpus/profile and rights; create a short text-origin book; audition two voices; review and repair a flagged passage; inspect regenerated navigation and hashes; show the commercial-audio path; navigate protected outputs on target hardware; inject/recover one worker failure; and review benchmark and handover evidence. Pre-generated complete books complement the live short example so the demonstration does not mistake a polished excerpt for long-form validation.

Understand the work

Scope and requirements

What the draft asks for, what it leaves open, and how to trace delivery evidence.

Download this article as MarkdownArticle permalink

The draft statement of work describes software for producing NLS digital talking books through two input routes. The engineering proposal is to share the downstream production system between those routes.

The draft background describes approximately 3,000 DTBs produced annually in the existing operation. That is context, not a numerical throughput acceptance target for the new system. SOW §1.

The two required input routes

Route Core work Important boundary
Existing commercial audiobook Ingest audio and metadata, identify divisions, create required supplemental speech, and construct the DTB. Proposed handling: preserve the supplied narration through approved mastering/conversion. A transcript supports analysis rather than replacing the recording.
Electronic text Ingest EPUB, Word, RTF, or plain text; recover structure; generate speech; construct the DTB. Preserve the source and document any change to its spoken form.

The commercial route explicitly accepts WAV, MP3, OPUS and FLAC, and requires AMR-WB+ conversion with MP3 encoding optional. The text route requires WAV and AMR-WB+ outputs, with MP3 encoding optional. These routes and formats come from SOW §§3.1–3.2. Scanned-book OCR is not an explicit core input in the draft. Treat it as a potential extension that needs scope agreement. A raw scan also creates recognition errors that the current text-first workflow would otherwise avoid.

Requirements and proposed evidence

The first two columns summarize the draft; the last is a proposed way to demonstrate completion, not a Government-approved acceptance checklist.

Area Draft requirement or concern Proposed evidence
Outputs NLS-compatible DTBs, navigation, metadata, required audio and related files; §§3.1–3.2 A reference book from each input route, validator results, and playback observations.
Processing Batch operation and automation with useful operator control Repeatable batch run, failed-job recovery, and an exception-review demonstration.
Integration Work within the Library's AWS environment and interfaces Deployment in the approved environment and a traced delivery to the agreed destination.
Security Federal security requirements referenced through Section H; §3.6 An agreed controls matrix after the missing section is supplied.
Compatibility and testing Library playback environment; §§3.11, 3.13 Test records for the agreed hardware, applications, and protected packages.
Ownership and continuity Non-subscription delivery and Library ownership; §§2, 3.12 Deliverable inventory, reproducible deployment, and dependency/voice-license records.
Documentation and support Technical documentation, training/support and warranty obligations Manuals, deployment runbook, handover exercise, support process and warranty terms.

Use the full SOW as the authority for exact wording; this table is a reading aid. The milestones page captures the concrete schedule terms.

Why local inference is the proposed default

The ownership and non-subscription language favors a system that can keep running without a required commercial speech API. The proposal is to deliver application code, approved model artifacts, configurations, and deployment machinery into the Library's environment.

That is an architectural interpretation, not a finding that every hosted service is forbidden or that open weights automatically satisfy the contract. AWS infrastructure still costs money. Third-party code, model weights, voice recordings, and commercial components each need their own rights assessment. Library ownership of delivered work does not erase third-party licenses.

Scope boundaries to settle

Agree which languages, book types, complex tables, equations, poetry, footnotes, and mixed-language passages are included in acceptance. Agree whether page navigation is required only when authoritative page references exist. Decide whether “various multimedia types” means an initial audio-plus-navigation product or additional full-text outputs. See open decisions.

A practical traceability record

For each accepted requirement, record its source section and edition, implementation component, test procedure, evidence artifact, and acceptance status. Keep this separate from model confidence scores. A model saying that a book looks correct is not proof of contractual acceptance.

Example proposal: the requirement for a valid navigation target maps to the compiler's reference checker, a deliberate broken-target test, and a playback test that reaches the expected chapter. This connects a rule to observable behavior.

Understand the work

Standards and output profiles

Resolve editions before treating DAISY or NLS compliance as one target.

Download this article as MarkdownArticle permalink

The supplied materials do not identify one unambiguous standards baseline. The draft repeatedly references Z39.86-2002, while the supplied ANSI/NISO attachment is the 2005 edition reaffirmed in 2012. Record and resolve this conflict before building an exporter.

Editions in the source package

Source Identified edition How to use it
Draft SOW ANSI/NISO Z39.86-2002; applicable NLS specifications, with current-version language Starting statement of requested work; request an explicit precedence ruling.
J1 construction specification NLS 1203:2018, version 1.3, January 2019 Attached construction rules, including legacy limits and exact file behavior.
J2 ANSI/NISO standard Z39.86-2005, reaffirmed 2012 Useful DTB specification, but not interchangeable with the 2002 baseline.
J3 navigation instructions QA201801, December 2018 Rules for interpreting and marking book structure.
Public NLS construction page 1203:2025 at review Comparison reference; not an automatic replacement for J1.

The public 1203:2025 specification still references the 2002 standard and removes the older 250-file limit. The attached edition and the public edition therefore need separate profiles until NLS confirms which applies. Public construction specification.

Why the edition changes code

An edition can change element names, allowed structures, metadata, counts, or validation behavior. For example, the 2005 NCX has pageList and pageTarget; the 2002 standard uses navList and navTarget for page navigation. A generic DAISY exporter can produce valid output for its own target and still miss the required NLS target. Compare the 2002 standard with the attached 2005 standard.

The SOW's reference to multiple multimedia types also needs clarification against J1's audioNCX metadata requirement (§5.3). The proposal is to deliver the approved audio-and-navigation profile first while preserving internal text structure for possible future outputs.

Proposed versioned profile

Treat an output profile as configuration with an immutable version, not scattered constants throughout the code. It should identify:

  • Governing standards and exact schema/DTD/entity files, with hashes.
  • Required metadata, navigation classes, announcements, and supported multimedia types.
  • Master-audio and compressed-audio parameters, timing rules, and container constraints.
  • File names, file-size/count limits, checksum format, protection rules, and ZIP behavior.
  • Approved validator versions and target playback systems.
  • Documented NLS clarifications or authorized deviations, including their scope.

Every build records its profile ID. A change in profile creates a new build lineage so that later reviewers can explain why two otherwise similar books differ.

A profile selection process

  1. Send NLS a short matrix of conflicting requirements, with source sections.
  2. Ask for written confirmation of editions, precedence, output types, and supplied tools.
  3. Turn the approved rules into a machine-readable profile and human-readable checklist.
  4. Compile a small reference book and test it with the approved tools and players.
  5. Freeze that profile for the pilot; handle later amendments as explicit version changes.

These are proposed implementation steps. The wiki does not establish legal or contractual precedence between conflicting documents.

Current references are separate from the baseline

The public mastering, protection, and delivery pages should be checked alongside the approved profile. Preserve dated copies or hashes when implementation begins; a live link may change after this wiki is written.

Design the system

System architecture

Separate AI interpretation from reproducible file construction.

Download this article as MarkdownArticle permalink

The proposed system has two front ends and one production core. Electronic text needs parsing and narration. Existing audiobooks need structure recovery and supplemental speech. Both converge on the same book representation, audio processing, compiler, validation, and delivery stages.

How the parts connect

Both input routes preserve their source evidence and feed the same book model, compiler and acceptance process.

Electronic text routeEPUB · Word · RTF · text Original text and source structure
parse and normalize
Source-linked book structure Preserve reading order, notes and page evidence.
Commercial audio routeAudio + available metadata Keep supplied narration through approved mastering.
transcribe, infer and align
Proposed navigation Use metadata, matching text and audio evidence.
converge on one versioned record
plan speech and retain existing audio
accept content and freeze the audio timeline
Deterministic constructionControlled by the approved output profile
XML + timing→Encoding→Protection→Packaging
inspect the actual package and playback
Independent acceptance Validators, target players and authorized review.
accept
Release and delivery Only accepted outputs reach the agreed NLS interface.
Proposed architecture. The input route changes preparation and speech generation; both routes use the same controlled construction and acceptance process. Failed checks return to the affected stage for repair.

Who does what

Work Proposed approach Reason
Recognize uncertain structure and ambiguous text Deterministic parsing first, model-assisted proposals when useful Source evidence stays available and uncertainty can be reviewed.
Generate narration and announcements Approved, replaceable speech models Voices and model families can be evaluated without rewriting the compiler.
Find possible spoken defects Independent ASR, alignment, acoustic checks, targeted review No single recognizer is a perfect judge.
Maintain IDs, source coverage, references and timing Typed data structures and conventional code These properties need exact, repeatable checks.
Construct XML, encode, protect and package Versioned deterministic compiler plus approved media tools Compatibility depends on exact rules and artifacts.
Accept a book Agreed validators, playback tests and authorized review Passing a model check is only one source of evidence.

The main modules

Ingestion stores the original files and parses format-specific structure. The book service manages immutable versions, source spans, metadata claims and approvals. Speech and analysis workers synthesize or inspect bounded segments. The compiler uses accepted content and a frozen timeline to emit the required files. Independent validators report defects. The review interface lets an operator resolve exceptions and trigger dependent rebuilds.

The delivery adapter publishes only accepted, complete outputs through the agreed NLS interface. It should not infer an interface from the fact that the environment uses AWS.

Separate three kinds of segmentation

A speech chunk is a unit that can be regenerated. A navigation element is a destination a reader can select. A delivered audio file is a packaging and playback unit. They do not need to have matching boundaries.

For illustration, one chapter could contain 80 synthesis chunks, 12 navigable sections, and one delivered chapter audio file. Those numbers are explanatory, not requirements. Keeping the concepts separate enables local repairs while respecting file limits and meaningful navigation.

Illustrative chapter · counts from the example below · not a timing scale

Three kinds of segmentation solve different problems. A correction can replace a small speech chunk while navigation and delivery remain organized at larger scales. The counts are illustrative, not NLS limits.

A book behaves like an incremental build

A pronunciation correction can affect several speech chunks. Regenerate those chunks, reassemble the affected masters, recalculate the dependent timeline, rebuild navigation/XML, encode and protect the changed content, then update checksums and packaging. Unaffected accepted synthesis can stay cached.

This is a dependency graph: each output records the exact inputs from which it was made. If an input changes, downstream evidence becomes stale. The build system should never keep a green validation result for an old version of a file.

Boundaries for AI assistance

Model outputs are proposals or generated media until verified. They must not silently delete source text, invent page numbers, substitute a narrator, write arbitrary final XML, or mark their own output accepted. Books and metadata are untrusted content, including instructions that appear inside them.

A useful model response is structured: a proposed chapter boundary, the source span supporting it, alternative interpretations, and confidence evidence. The compiler consumes an accepted decision, not a free-form conversation.

Small services before many services

The module boundaries can initially live in one application with separate worker processes. Splitting every module into a network service is not necessary to prove the approach. Establish stable data contracts, job IDs and reproducible artifacts first; deploy independently where isolation or scale justifies it.

Design the system

Canonical book model and rebuilds

Keep source passages, spoken text, audio, navigation and evidence connected.

Download this article as MarkdownArticle permalink

The canonical book model is the system's shared record of what the book contains and how each part becomes audio and navigation. “Canonical” means that downstream stages refer to the same structured representation; it does not mean that an AI transcript replaces the original source.

Proposed invariant: every source passage is represented in the output or has an explicit, reviewable reason for exclusion. This is the strongest protection against quiet omission, duplication and paraphrasing.

Follow one passage through the system

Proposed data model · trace one passage through a build
Source passageOriginal text + stable location$1,205
→
Approved spoken text“one thousand two hundred five dollars”
→
Accepted audioVersioned asset + exact sample positions

First arrow: retain the normalization rule and approval. Second arrow: retain the model, voice and lexicon versions.

Source span + stable node IDPreserves hierarchy and reading order
→
Structure recordHeading, paragraph, note, table or page marker; parent + order
Structure + accepted audioResolves logical destinations into media
→
Two navigation referencesMain narration destination and separate spoken heading-label range
Approved text + observed audioASR is evidence of what a recognizer heard
→
Quality recordArtifact version, check, evidence, defect and resolution; ASR stays separate from approved text
Exact input versions + dependent checks
Auditable output buildProfile and tool versions, dependency records and package hashes identify exactly what was delivered.
One passage, connected records. This proposed model preserves the source while linking each transformation, media asset and approval to an exact version. The currency example illustrates normalization; it is not an NLS-prescribed phrase.

What to retain

Record Essential information Why it matters
Source artifact Original bytes, hash, format, ingestion version Reproduce and audit a conversion.
Source span File and stable structural location; original text or audio range Locate the evidence behind every transformation.
Structure node Heading, paragraph, note, list, table, page marker; parent and order Preserve reading order and navigation semantics.
Text transformation Original text, approved spoken form, rule, rationale, reviewer Distinguish faithful normalization from accidental rewriting.
Speech recipe Model/weights version, voice asset, lexicon version, settings Explain and selectively regenerate the audio.
Audio mapping Asset hash, sample rate, start/end sample positions, source span IDs Tie content and navigation to actual media.
Quality record Check type, artifact version, evidence, defect status, resolution Prevent stale approvals from surviving changes.
Output build Profile version, tool versions, dependencies, package hashes Explain exactly what was delivered.

This schema is a proposal. Its final field names and storage technology are implementation choices.

Preserve three forms of text

Store original written text, approved spoken text, and observed ASR text separately. Each answers a different question: what did the source say, what should be spoken, and what did a recognizer hear?

Example:

Original text:  $1,205
Spoken text:    one thousand two hundred five dollars
Transformation: currency/USD
Source location: chapter 4, paragraph 12, characters 18–24
ASR evidence:   populated only after examining the generated audio

A normalization rule should never convert “fifty” and “fifteen” into the same comparison token. Preserve high-consequence details such as negation, units, dates and proper names.

Use stable IDs and immutable versions

Give structural nodes stable identities across edits when they represent the same passage. Give source snapshots and generated assets immutable version IDs or content hashes. A node can point to a new approved audio asset without overwriting the old one.

A timestamp alone is an inadequate identity: earlier edits can shift every later timestamp. Store the logical destination and the resolved audio position, then derive the latter for each completed build.

Track two independent audio references

A navigation item points to a destination in the main narration and may also reference a spoken label in a separate headings file. Record both asset identities and timing ranges explicitly. Confusing the two timelines can produce a correct-sounding label that jumps to the wrong place.

Explore which work changes

Proposed dependency model · explore a change

A lexicon correction regenerates affected speech. Downstream timing and packages follow the changed audio; unrelated accepted synthesis can be reused.

Rebuild or rerunReuseDepends on scope
Original source + structureReuseRetain the source snapshot and stable passage identities.
Main narration chunksRebuild affectedRegenerate passages using the corrected lexicon entry; reuse unrelated accepted chunks.
Spoken heading labelsDepends on scopeRegenerate only labels that use the corrected entry.
Assembled audio + final timingRebuild affectedReassemble affected masters and calculate final sample positions.
Navigation + XMLRebuild dependentRebuild dependent SMIL/NCX references and durations from the new timeline.
Encode, protect + packageRebuild changedProcess changed media and update dependent protection, checksums and packages.
Validation evidenceRerun dependentValidate the changed dependencies and completed build; prior approvals remain tied to prior hashes.

Scope matters: a surname can have different pronunciations in different contexts. Record which occurrences the accepted correction covers.

A local edit can reach the final package. This simplified dependency order illustrates the proposed rebuild strategy. Narration and heading labels have separate assets and timing references. The approved profile determines whether a heading clip is copied or synthesized. The implementation must follow the actual dependency graph.

Worked repair example

  1. An operator corrects a surname in the book lexicon.
  2. Dependency lookup finds the passages whose spoken forms use that entry.
  3. Those passages receive new speech recipes and audio versions.
  4. The affected assembled masters and their final sample positions change.
  5. Dependent SMIL/NCX timing, durations, checksums, protection and packages are rebuilt.
  6. Validation runs against the new artifact hashes; the prior build remains auditable.

The system should record who accepted the correction, why it was made, and whether all occurrences should change. A name may intentionally have different pronunciations in different contexts.

Define cache identity carefully

Include source/spoken text, model and voice artifacts, normalization and lexicon versions, generation parameters, and relevant processing versions. Store the accepted audio itself. A random seed cannot guarantee byte-identical future synthesis across different hardware or software versions.

Design the system

Source ingestion and navigation recovery

Preserve existing structure and make uncertain structure reviewable.

Download this article as MarkdownArticle permalink

Ingestion should retain source evidence before attempting to improve it. The workflow differs by input format, but every importer should produce source-linked content and proposed navigation for the common book model.

Electronic text formats

Format Recover first Watch for
EPUB Package metadata, reading order, XHTML structure, headings, links, notes, authoritative page mapping Visual styling that disguises semantic structure; broken reading order.
Word Heading styles, outline levels, lists, tables, footnotes and endnotes Direct formatting that looks like a heading but has no heading style.
RTF or legacy Word Convert in an isolated process while keeping the original and a conversion record Lost notes, tables, symbols, or outline information.
Plain text Delimiters and deterministic heuristics, followed by model-assisted classification Missing hierarchy and ambiguous headings; uncertain reading order.

The format list comes from the SOW §§3.1–3.2; the importer design is proposed. Do not flatten a structured source into plain text before extracting its relationships.

Main content determines navigation

J3 §3.1 directs navigation work to the main content rather than simply copying the table of contents. A table of contents may omit subheadings or describe them differently. Inspect the actual text and preserve evidence for every proposed division.

J3 also separates hierarchical navigation from nonhierarchical structures such as notes and segments. A note is not necessarily another chapter. The correct distinction affects playback behavior and must survive ingestion into the compiler.

A model can propose that a passage is a section heading, but the proposal should include its source location and supporting evidence. Printed page numbers should come from an authoritative page map or source, not model invention.

Commercial recordings

Use evidence in this proposed order: publisher metadata and chapter boundaries, matching electronic text when available, an ASR transcript, semantic classification, and precise audio alignment. Record conflicts rather than silently choosing whichever source arrived last.

For example, a model can distinguish a spoken chapter announcement from dialogue mentioning “chapter three.” The final destination still needs alignment and a waveform-level boundary check.

Keep the supplied narration through approved mastering and conversion. The ASR transcript is an analysis aid that can contain errors. Synthesize only the additional speech the approved output requires. Availability of a narrator's recording is not permission to clone the narrator's voice.

A recording cannot reliably reveal an unspoken print heading or an unannounced print page number. Mark these as unavailable, and agree how that input class will be accepted.

Written text to spoken text

Normalization expands abbreviations, dates, numbers and symbols while retaining original source spans. The sentence “Dr. Smith paid $1,205 on 03/04/2026” needs context: “Dr.” may require disambiguation, the currency needs expansion, and the date convention is ambiguous.

Use deterministic rules for well-defined cases and a reviewable proposal for contextual cases. A book-level pronunciation lexicon keeps repeated terms consistent. NVIDIA NeMo text processing is a candidate foundation for text normalization, not a guarantee that the whole book is correctly interpreted.

Model adapters and source safety

Each speech adapter should declare which controls it supports: SSML, phonemes, pronunciation dictionaries, language selection, reference audio, and style instructions. Unsupported controls must fail clearly or use an agreed conversion; they should not be silently ignored.

Treat model-specific control strings in source text as content. Escape or otherwise neutralize them before a model can interpret them as a voice or performance instruction. Retain the literal source and record the transformation.

What an importer should report

Proposed outputs include an ordered content tree, source coverage map, metadata conflicts, lost-format warnings, navigation proposals and unresolved passages. A successful file parse alone is insufficient: compare recovered content and reading order with the source before moving into speech production.

Design the system

Speech generation and model evaluation

Compare candidates on full-book fidelity and repair effort.

Download this article as MarkdownArticle permalink

Use a replaceable text-to-speech (TTS) interface and evaluate several local models on the same material. The names below are candidates from the original analysis. None has been demonstrated here to satisfy NLS acceptance requirements.

Prototype model selection

The prototype model plan selects Qwen3-TTS 1.7B CustomVoice as the initial narration candidate, with Kokoro as a comparison baseline. It narrows the original broad shortlist into explicit Mac/CUDA runtimes, supporting models and evaluation conditions.

Candidate roles

Candidate Documented characteristics Proposed role
Qwen3 TTS Local inference and multiple model/voice modes; the named 1.7B CustomVoice weights list Apache 2.0 Proposed first adapter uses CustomVoice presets; Base reference-voice workflows are conditional. See the prototype model plan.
Chatterbox Official project lists an English Turbo model and multilingual options under an MIT project license Independent challenger for consistency and difficult passages.
Kokoro 82M Small open-weight TTS model; card lists Apache 2.0 Compact baseline for narration and announcements.
Fish Audio S2 Pro Model-card licensing requires attention, including separate commercial licensing Conditional evaluation candidate after terms are resolved.

License labels do not establish rights to every associated voice, dataset, dependency or deployment use. Pin the exact artifacts and record applicable terms in the handover inventory. This wiki's model descriptions are not a legal license clearance.

A book level narration plan

Before synthesis, select an approved voice, pronunciation lexicon, language handling, and restrained pacing/style settings. Preserve these across chunks. Faithful reading is the default: selectable voices do not imply automatic character acting, added laughs, paraphrasing or sound effects.

Test paragraph or sentence-group chunks. The original analysis suggests experimenting with 20–60 seconds of speech per chunk; this is a tuning range, not a requirement or measured optimum. Preserve sentence boundaries, and compare voice continuity, pronunciation, joins, and correction cost at different lengths.

Bounded generation and repair

  1. Synthesize a chunk from approved spoken text and a recorded recipe.
  2. Transcribe with an independent recognizer and compare meaningful content.
  3. Check alignment, clipping, silence, level and joins.
  4. Repair a specific discrepancy through a lexicon, text, boundary or generation change.
  5. Escalate after an agreed attempt limit rather than retry indefinitely.

Do not silently switch narrators for a difficult sentence. An engine fallback must preserve the approved voice or trigger a deliberate review and potentially a larger regeneration.

ASR and alignment candidates

Qwen3 ASR provides transcription and a separate forced aligner. WhisperX combines transcription with alignment and documents limitations. Long books need bounded jobs with reconciled overlap, rather than a single whole-book alignment call.

ASR asks what a recognizer hears. Forced alignment asks where expected text fits the audio. These are different checks, and neither establishes perfect narration alone. See the validation page.

A useful evaluation corpus

Include ordinary prose and long chapters alongside names, dates, quantities, negation, abbreviations, dialogue, notes, lists, tables, mixed-language passages, and repeated terms. Add passages from each supported input class. Keep tuning material separate from a held-out acceptance set.

Measure complete books as well as excerpts. Report unresolved defects per accepted audio hour, review minutes, repair attempts, throughput, and consistency across joins. Do not infer whole-book quality from a 30-second demo or time to first audio.

Voice governance

Keep approved voice assets, consent or usage rights, version, language scope and allowed transformations with the production recipe. A change to the voice reference can affect every generated passage and therefore should invalidate the appropriate cache and quality evidence.

Design the system

DTB compilation and packaging

Turn accepted content into exact, validated navigation and media files.

Download this article as MarkdownArticle permalink

The compiler converts an accepted book model and finalized audio into the files required by an approved NLS output profile. It should serialize typed data and enforce explicit rules. Generating final XML directly with an LLM would make exact compliance harder to control.

What the main files do

File or format Role
OPF Describes metadata, the package contents and reading/playback organization.
NCX Defines navigable structures, labels and destinations.
SMIL Connects ordered playback structures to particular media clips.
WAV masters Accepted uncompressed source audio used as the timing reference.
AMR-WB+ in 3GP Compressed audio with NLS-specific codec and container constraints.
Resource files Supply approved alternative navigation-class labels when needed.
dtb.md5 XML checksum document for the DTB files, excluding itself.
DTD and entity files Local dependencies needed to interpret and validate the XML.

These descriptions summarize J1 §§3.1–3.10. Exact applicability depends on the selected profile.

Freeze audio before final timing

Proposal: track positions as integer samples, complete approved mastering and assembly, then freeze the accepted WAV timeline. Derive clip timestamps from that final timeline, not from rough ASR estimates or an earlier synthesis chunk.

In attached J1 §3.2.5, timestamps remain relative to the accepted WAV; codec offset compensation is assigned to the player. Do not silently shift XML to encoded-audio timing or compensate twice. J1 pp10–11.

Timing windows depend on clip type

Hatched band = permitted location of the clip boundary
Read from the speech edges outward. Each zero marks the adjacent speech boundary. Lead-in and tail use separate scales; the speech block has no time scale. Select one endpoint within each permitted band. These are boundary windows, not required amounts of inserted silence. If a valid boundary is unavailable, flag the case or approve an audio edit and rebuild timing.
Clip type Start before speech End after speech Attachment source
Main SMIL narration clips 80–120 ms 150–300 ms J1 §3.3.4, p11
Heading labels 80–120 ms 150–300 ms J1 §3.4.5, p15
Alternative-class resource audio 100–200 ms 150–500 ms J3 §3.3.3, p8

The first two rows also require the boundaries to lie in silence. Both clip endpoints must be present. A shared global “navigation padding” constant would miss the resource exception.

Use alignment plus local waveform analysis to identify the actual speech boundary. If the required margin is unavailable, flag the case or make an approved audio edit and rebuild dependent timing. Choosing the middle of a permitted window is a proposed default, not an NLS rule.

Main destinations and spoken labels

A navigation destination resolves into primary audio through SMIL. Its spoken label may refer to a different range in the headings audio. Keep both references explicit. Labels must represent the heading, not the whole section. J1 §§3.4.3.4–3.4.3.5, pp13–14; J3 §3.1.2, p6.

J3 requires at least two level-one navPoints: the first marks the opening title/author with class title/author, and the last marks the closing announcement with class close. J3 §3.1.1, p5.

Notes, note references, sidebars, pages, line numbers and segments need their own semantics. J3 requires segment navigation for entries in specified lists, and J1 says not to create navLists for segments. Treating every navigation object as a chapter loses these distinctions. J3 §§3.2.1–3.2.4, pp6–7; J1 §3.4.4, p14.

Exact rules from the attached construction edition

Rule in J1 Compiler consequence
No more than 250 files; §3.1.3 Count all required files under that edition. The newer public edition differs.
Split SMIL above 100 KiB; at most 50 SMIL files; §3.3.12 Measure actual UTF-8 serialized bytes; 100 KiB is 102,400 bytes. Follow the stated escalation rule if limits are exceeded.
At most 5,000 navPoints; contact monitor when exceeding 1,000; §3.4.3.6 Distinguish a hard limit from a consultation trigger.
Lowercase filenames and prescribed production IDs/suffixes; §3.1 Generate names from the profile and NLS-supplied identity.
XML diskcheck checksum structure; §3.9 Ordinary md5sum text is insufficient. Include each file except the checksum itself; do not list the checksum in the OPF manifest.
Include referenced DTDs and entities; UTF-8 XML; §3.10 Package dependencies and validate against an approved local catalog.
Duration follows SMIL playback, including specified custom-test content; §5.3(v) Derive duration from the compiled playback structure, within the stated ±1 second accuracy.

See J1 pp9–20. Do not apply the attached 250-file limit to a newer profile by accident. Permit only approved local XML dependencies; do not enable arbitrary network/entity resolution for book content.

Encoding is an early interoperability risk

AMR-WB+ is not ordinary AMR-WB. Attached J1 specifies mono constant bitrate, frame type 23 and ISF index 8, plus particular 3GP metadata and field constraints. A file extension alone proves none of those properties. J1 §3.2.3, p10.

Hindenburg's support article describes an NLS-supplied encoder; APH's revision history records NLS protection/validation integration. These are historical precedents. Obtain current components, terms, supported operating systems, automation interfaces and cloud-deployment approval before relying on them.

Protection is part of the book

PDTB protection involves protected content and associated authorization/key material. S3 storage encryption protects stored objects but does not create a protected talking book that an NLS player can authorize. See the DAISY protection specification and request the governing NLS 1205 profile and tooling.

Packaging and final submission

The public 1206:2025 delivery specification describes an outer ZIP with protected/unprotected DTB archives and WAV masters. It specifies uncompressed entries and local-header CRC/size values with ZIP bit 3 clear. A seekable staging file is a proposed way to control those headers, followed by archive inspection and extracted-file checks. Delivery specification §§3–3.2.

Its §4 names the NLS Transfer Portal. Internal S3 staging is a separate architectural choice; confirm the project-specific final interface. Keep QA logs outside the delivery package unless that interface expressly permits them.

Reuse existing DAISY components selectively

DAISY Pipeline's DTBook-to-DAISY3 script accepts a 2005-3 DTBook input and offers speech-related conversion options. It is a useful candidate for selected transformations. That documentation does not establish an NLS 2002, AMR-WB+, protected export path. Keep the NLS profile and acceptance tests explicit even when reusing its code.

Prove and operate

Quality checks and operator review

Detect defects, repair bounded passages, and preserve acceptance evidence.

Download this article as MarkdownArticle permalink

Automation succeeds when defects are detectable and repairs are precise. A generated file is only an intermediate result until content, audio, navigation, format and playback checks agree with the accepted requirements.

The following layered QA design is proposed. The draft requires quality and compatibility testing but does not provide numerical narration-error or review-effort thresholds. Those must be agreed before the pilot is judged. SOW §§3.5, 3.11, Phase II.

From a candidate to an accepted release

Proposed control flow. The retry budget, acceptance criteria and review authority must be agreed. The draft calls for quality and compatibility testing; the numerical acceptance thresholds still need agreement. Draft SOW §§3.5, 3.11 and Phase II.

Six layers of correctness

Layer Example defect Proposed evidence
Source fidelity A paragraph is omitted or reordered Source-span coverage and ordering checks against the imported book.
Speech fidelity A number changes or a word is repeated Independent transcription, protected-token comparison and targeted listening.
Acoustic quality Clipping, an abrupt join, long silence or voice drift Signal checks plus excerpt and sustained-listening review.
Navigation A label announces one chapter and lands in another Reference resolution, alignment and interaction tests.
Format Broken XML references, wrong codec, bad checksum or ZIP headers Schema, semantic, media and package inspectors.
Playback Protection, skip, resume or speed behavior fails on a target player Recorded test matrix using actual agreed players/apps.

ASR evidence is not ground truth

An ASR mismatch can arise from a TTS mistake, a recognizer mistake, or a legitimate difference between written and spoken forms. Forced alignment fits expected text to an audio signal; it cannot independently prove every expected word was spoken.

Combine independent transcription, alignment, acoustic evidence and targeted checks for names, quantities, units, dates and negation. For high-consequence discrepancies, use another recognizer or a human listener. Agreement between models increases evidence but does not guarantee correctness. WhisperX's documented limitations illustrate missing timestamps and other alignment problems.

A review screen should show enough evidence

For each issue, present the source passage, approved spoken text, recognizer output, short playable excerpt, timeline/waveform, reason for the flag, affected build and repair history. Offer concrete actions: correct pronunciation, edit a proposed boundary, regenerate a passage, or accept a documented exception.

An exception should retain rationale and authorized reviewer identity. Acceptance of one passage should not automatically suppress the same error class throughout a book.

Defect states and bounded retries

A proposed workflow is detected → triaged → repair attempted → rechecked → accepted or escalated. Set an agreed retry budget per issue. Distinguish processing failures, which may be safely retried, from content ambiguities, which often need a decision.

Retest everything invalidated by a repair. Changing a chunk's duration can invalidate later clip times even when the later audio remains unchanged. Tie every check result to the exact artifact hashes and output profile.

What to measure

Metric Definition to agree Why it helps
Unresolved content defects Confirmed omissions, substitutions or repetitions per accepted audio hour Tracks fidelity more directly than general “naturalness.”
Navigation accuracy Correct labels and destinations divided by evaluated items, by type Prevents chapter success from masking note/page failures.
Review effort Active human review minutes per accepted audio hour Measures the practical cost of automation.
Repair burden Attempts and regenerated audio per accepted hour Exposes unstable models and costly chunk sizes.
Throughput Accepted output hours per elapsed production hour and per compute hour Separates machine speed from accepted-book delivery.
Operational reliability Failed/resumed jobs, incomplete packages and recovery time Tests production behavior rather than isolated demos.

These are proposed definitions, not measured results or contractual targets. Report the test corpus, input class, language, hardware, model version and accepted-output denominator with every result. Avoid hiding rejected books from quality statistics; report rejections separately.

Playback tests beyond the opening chapter

Exercise all navigation levels, notes/sidebars and skipping behavior, segment stepping, resume, transitions between audio files, and faster playback. Check both protected and unprotected artifacts where relevant. Record the player/app version and profile.

The draft expressly calls for testing with the Library playback environment. The exact device matrix and pass criteria remain open decisions. Accessibility of the operator interface should also be tested with keyboard and assistive technology as a proposed implementation practice.

Keep acceptance independent

Use known reference books, deliberate malformed examples, and a held-out evaluation set. An AI coding assistant can help generate tests and analyze regressions, but the acceptance criteria and reference outputs should be reviewed independently from the implementation that produces them.

Prove and operate

AWS deployment and production operations

A proposed deployment, realistic sizing method, and operational handover.

Download this article as MarkdownArticle permalink

The draft requires integration with the Library's AWS systems and standard import/export without proprietary lock-in. The components below are a proposed implementation, not a mandated stack. Confirm available services, network boundaries, security rules and delivery interfaces with the Library. SOW §3.7.

Prototype hardware and cloud boundary

See Mac hardware and cloud options for proposed development tiers, single-GPU AWS workers and an adjustable capacity worksheet. The supplied draft supports AWS integration; it does not provide an allowlist for other providers or settle account, region or service approval.

How deployment responsibilities connect

Operator / reviewerStarts jobs, resolves exceptions, authorizes release
uses the review application
Proposed components in the Library’s approved AWS environment
Application + control planeBook versions, jobs, approvals and output profiles
state
PostgreSQLStores state and decisions for the control plane
schedules jobs and records results
Workflow orchestrationStep Functions + AWS Batch, or an approved equivalent
dispatches bounded, resumable work
Model workersTTS · transcription · alignment
Pinned models and voices; GPU or suitable CPU
Compilation workersParsing · XML · media tools · validation · packaging
Isolated CPU jobs
workers read inputs and write versioned artifacts
S3 artifact storage
Originals→Intermediates→Staging→Accepted outputs
Delivery adapter
Submits only the accepted package; records acknowledgement
crosses the agreed delivery interface
NLS receiving workflowPublic 1206 guidance names the Transfer Portal.
The project-specific integration still needs confirmation.
Proposed deployment, not a prescribed AWS stack. The control plane manages decisions, workers process content, S3 stores artifacts, and a separate delivery adapter submits accepted books. S3 storage is not itself the confirmed submission interface. The diagram shows logical responsibilities, not a network-security design.

Proposed deployment components

Component Candidate implementation Responsibility
Control plane Typed application service Book versions, jobs, profiles, approvals and audit events.
Metadata and job state PostgreSQL Structured records, transactions and searchable status.
Artifact storage S3 Separate originals, intermediates, staging and accepted outputs.
Speech and analysis workers Containerized Python workers with pinned models GPU inference or appropriate CPU inference.
Compilation workers Isolated CPU containers Parsing, XML, media inspection and packaging.
Workflow orchestration Step Functions with AWS Batch, or approved equivalent Dependencies, bounded retries, checkpoints and review gates.
Operator application Accessible web interface Exception review, progress and authorized release.
Delivery adapter Agreed NLS integration Submit validated packages and record acknowledgement.

AWS documents GPU jobs in Batch and Step Functions integration with Batch. These capabilities support the design, but do not establish regional availability, account quotas, approval or required capacity.

Jobs should be repeatable and resumable

Use immutable inputs, explicit output versions and idempotent job keys. A retried job should not duplicate a submission or overwrite an accepted book. Record checkpoints after durable artifacts are written and verified. Use separate states for failed processing, unresolved content, ready for review, and ready for delivery.

Keep model workers warm over a batch where appropriate. Loading weights for every paragraph wastes time, but the right worker lifetime depends on measured throughput and utilization.

Capacity without invented benchmarks

A useful planning equation is:

GPU-hours of work = source audio-hours × processing factor × retry multiplier
Wall-clock hours ≈ GPU-hours of work ÷ (GPU count × effective utilization)

Here the processing factor must be measured in GPU-hours per source audio-hour for the chosen workload. If it is a combined factor, avoid counting transcription or repair twice.

For illustration only: 30,000 hours × 0.5 GPU-hours per hour = 15,000 GPU-hours before retries. At a 1.2 retry multiplier, this becomes 18,000. Ten GPUs at 75% effective utilization would imply about 2,400 elapsed hours of compute processing. These are assumptions and arithmetic, not observed performance or a calendar delivery forecast. Human review, CPU work, queues and delivery can remain bottlenecks.

Benchmark the two input paths separately. Commercial recordings may need full-book transcription/alignment but only supplemental TTS; electronic text needs full narration. Measure accepted output, including failed attempts, rather than quoting isolated generation speed.

Security and source handling

The draft references federal controls in Section H, which is absent from the supplied attachments. It also calls for source/output protection, retention/destruction procedures and security documentation. Obtain that missing material before fixing the controls baseline. SOW §§3.6, 3.13.1, 5.1.

Proposed defaults include isolated inference in the approved environment, restricted unnecessary egress, pinned dependencies, least-privilege roles, separate content-protection keys, and limited source text in logs. Book content must be treated as data, including text that looks like instructions. SOW §5.1 limits use of Government source materials to contract performance and requires secure disposal at contract end. Implement those obligations alongside the agreed retention and deletion procedures.

Storage encryption, transfer encryption and PDTB protection solve different problems. Keep their responsibilities and key handling explicit.

Handover and maintenance

Deliver source code, deployment instructions, configuration, dependency/model/voice inventory, operator and administrator guides, recovery procedures, validation tools and a known reference build. Confirm which third-party components can be redistributed and which must be supplied by NLS.

The draft specifies support response within 24 hours and a 12-month defect-correction warranty after acceptance; it permits ongoing support/maintenance separately from prohibited recurring access requirements. See the milestone page and SOW §§3.8–3.9, 3.12.

Release evidence

For each released software/profile version, preserve migration notes, a reproducible deployment recipe, dependency hashes, regression results and the reference-book outputs. That makes later maintenance possible without relying on the original contractor's environment.

Prove and operate

Proof gates and draft milestones

Separate the proposed technical sequence from the actual SOW schedule.

Download this article as MarkdownArticle permalink

Use technical proof gates to reduce uncertainty while meeting the draft's separate contractual phases. The five gates below come from the engineering proposal; they are not the four phases named in the SOW.

Detailed prototype work plan

The prototype implementation plan adds a proposed 16-week engineering sequence, scope boundaries, owner roles and handover artifacts. The acceptance plan defines the proposed corpus and evidence. These planning estimates do not replace the draft deadlines below.

Proposed proof gates

Gate Work to demonstrate Pass evidence
1. Format and protection Build a tiny complete book with the approved profile, encoder, protection and packaging Approved validation results and working navigation/playback on target systems.
2. Long-form speech Compare candidate engines on chapters and complete representative books Agreed fidelity and quality results, with measured correction effort.
3. Both ingestion paths Process structured text, less-structured text, commercial audio with and without matching text Source coverage and independently checked navigation, including explicit unavailable information.
4. Repair and rebuilds Change pronunciation, replace a paragraph, edit a heading, interrupt processing and fail encoding Only necessary work repeats; all dependent timing, hashes, metadata and packaging remain correct.
5. Production operation Run batches, security checks, recovery and real integration Repeatable accepted outputs, measured throughput and operational handover.

Prove codec/protection compatibility early with simple known content so that format errors can be separated from speech errors. Speech and importer experiments can proceed in parallel after their required inputs are available.

Schedule in the supplied draft

POP means period of performance. No actual start date or total duration is supplied here, so the dates below are relative triggers rather than calendar commitments. SOW §4.1 and named phase descriptions.

Deliverable or event Draft timing
Key personnel resumes At proposal submission.
Kickoff Within 10 business days; the table uses POP start while §5.2 uses contract award. Clarify the trigger.
Five architecture diagrams and eight technical documents Within 30 calendar days after POP start.
Status and quality-control reports Every 30 calendar days during POP.
Phase I prototype, sample DTBs, workflow demonstration and test report Within six months of POP start.
Phase II integration and operational testing Starts after Phase I completion and Government acceptance; completed within 90 calendar days.
Phase III full deployment Starts after Phase II completion and Government acceptance; completed within 90 calendar days.
Phase IV ongoing support After Phase III completion and acceptance, through the remainder of POP.
Warranty 12 months after written Government acceptance of all deliverables.

The 24-hour support-response requirement appears in §3.8. The warranty includes free defect correction and patches under §3.9. Do not turn the phase durations into a promised total project length: Government acceptance timing and the total POP are not established here.

How the gates can fit the phases

Proposed mapping: use Phase I to resolve the profile and prove format/protection, model behavior and initial ingestion. Continue hardening ingestion, repair and operational behavior through Phase II integration tests. Phase III adds deployment, handover and acceptance evidence. Phase IV covers production support and controlled improvement.

This mapping does not reduce the deliverables required in any phase. Validate it against the final solicitation and approved project plan.

Early documents should expose real decisions

The first-month architecture package should show both source paths, data ownership, security boundaries, failure/retry handling, and the split between AI proposals and accepted compiler inputs. List unresolved standards and supplied-component dependencies rather than hiding them behind a generic cloud diagram.

Suggested readiness check before a pilot

Confirm the output profile; acquire reference books, target players and required tooling; agree supported content classes and quantitative acceptance criteria; record authorized voices and dependency rights; obtain Section H; and identify who can accept exceptions and release a book. These are suggested prerequisites tied to the decision register.

Understand the work

Open decisions and questions for NLS

A concrete register of answers needed before committing to implementation.

Download this article as MarkdownArticle permalink

The items below remain unresolved by the supplied materials. Suggested decision partners describe the kind of authority needed; they are not assignments made by NLS. Do not mark an item resolved without a written answer, accepted test result or other attributable evidence.

Decisions that affect the architecture

Question Why it matters Evidence needed and suggested decision partner
Which editions and precedence govern the 2002/2005 conflict and current-version language? Changes XML structures, limits and test rules. Approved profile matrix from the contracting/technical authority.
Which multimedia types are required at acceptance? audioNCX and synchronized full text imply different outputs. Supported output-type list and reference books from NLS production/QA.
Which encoder, protector, validators and keys will NLS provide? Determines the feasible encoding/protection route and deployment platform. Current packages, versions, terms, headless interfaces, OS support and cloud approval.
What accompanies commercial recordings? Metadata and matching text determine recoverable structure and pages. Representative source sets and expected navigation from NLS production.
How do mastering rules apply to synthetic audio and generated labels? Native model rate, initial-recording language and copied heading rules may conflict with a chosen workflow. Approved interpretation and accepted synthetic example from NLS audio/QA.
What security controls and retention rules apply? Determines network, access, key, logging and deletion design. Missing Section H and approved architecture requirements from Library security.
What constitutes an accepted book? Controls quality thresholds, human review and pricing/effort assumptions. Content classes, language scope, player matrix, defect thresholds and throughput/turnaround targets.
What is the final submission interface for this project? The public Transfer Portal and proposed internal S3 staging are different interfaces. Integration contract, credentials/process ownership and acknowledgement semantics.
How are third-party model, voice and component rights handled? Library ownership/non-subscription terms must coexist with component rights. Approved inventory and applicable license/voice permissions from appropriate procurement/legal reviewers.

Specific mastering clarification

The live public mastering page is 1202:2026, revision 1.5, June 3, 2026. It specifies initial 44.1 kHz PCM and adds integrated loudness of −22 LUFS ±2; it also says heading clips are copied from primary audio. The public index still advertises a 2025 entry. Confirm the actual governing edition and how these rules apply to neural synthesis and separately generated navigation speech. Resampling a model's output alone does not settle the “initially recorded” interpretation. Current mastering page and specification index.

Additional source gaps

The SOW mentions 1201; J1 references 1208 for commercial-audio masters; only 1203 is among the four attachments. Request the complete applicable specification set. The phrase “and D standards” in SOW §3 is unexplained. The kickoff start trigger differs between §4.1 and §5.2. These are useful clarification items, not grounds to invent a preferred interpretation.

Proposed decision record

For each answer, store the exact question, governing source and section, answer date, responsible authority, evidence link, affected profile/version, affected tests and remaining conditions. Preserve superseded decisions as history.

Example: approving a new file-count rule updates the output profile and the corresponding validator. It should not quietly change an already submitted book or erase the rule used for its original build.

What can proceed while answers are pending

Build the source-preserving data model, importer experiments, independent quality checks and model evaluation harness. Prototype against a clearly labeled provisional profile. Avoid committing to a production codec/protection stack or promising numerical quality/throughput until the relevant components and acceptance rules are confirmed.

Reference

Talking book glossary

Plain-language definitions for the terms used throughout the wiki.

Download this article as MarkdownArticle permalink

This glossary explains the terms as used in this proposed production system. The linked articles and standards contain the implementation detail.

Term Meaning
NLS The Library of Congress National Library Service for the Blind and Print Disabled.
DTB Digital talking book: audio and supporting structure that allow accessible reading and navigation.
DAISY Standards and technology for accessible publications, including structured talking books.
ANSI/NISO Z39.86 The talking-book standard cited in the procurement; the edition matters.
audioNCX A multimedia type describing full audio with navigation structure; it does not by itself promise a synchronized full-text product.
TTS Text to speech: generating spoken audio from text.
ASR Automatic speech recognition: generating a transcript from audio.
Forced alignment Estimating where a supplied transcript fits the audio; distinct from independently recognizing what was spoken.
Canonical book model The shared structured record linking original sources, approved text, navigation, audio and evidence.
Provenance Records of where an item came from and how it was transformed.
Source span An identifiable passage or audio range in an original source.
Spoken form The approved text representation intended for narration, such as expanding a currency amount.
Lexicon A book-level or shared set of pronunciation decisions.
Inference Running a trained model to generate or analyze content.
Model weights Numerical parameters of a trained model; they are distinct from application code and voice assets.
Deterministic compiler Conventional software that produces prescribed files from explicit accepted inputs and a versioned rule set.
Output profile The exact standards, media, metadata and packaging rules chosen for a build.
OPF The package document describing metadata, included resources and organization.
NCX Navigation Control file for XML; defines navigation structures and labels.
SMIL Synchronized Multimedia Integration Language; links playback structures to media segments.
DTD Document Type Definition; a schema mechanism used by these XML formats.
Entity An XML declaration/reference whose allowed definitions must be controlled and packaged as required.
navPoint A hierarchical navigation point, such as a chapter.
navTarget A target in an auxiliary navigation list; its allowed use depends on the standard edition.
Segment A useful navigable unit that need not be a chapter or heading.
PCM / WAV PCM represents uncompressed audio samples; WAV is a common container for those samples.
Sample position An integer location in audio at a known sample rate, useful for precise internal timing.
AMR-WB+ The extended wideband audio codec required by the cited NLS profile; different from AMR-WB.
3GP The container used for the compressed audio, with NLS-specific constraints.
LUFS A unit for measured perceived loudness; the applicable mastering edition defines the target.
PDTB Protected Digital Talking Book; content protection and authorization for the book/player ecosystem.
MD5 A checksum algorithm used in the specified DTB checksum format. A checksum is not an authorization or authenticity system.
CRC A ZIP integrity check value; the delivery specification constrains how it appears in the headers.
Idempotent job A job whose retry does not create unwanted duplicate effects.
Incremental build Regenerate changed outputs and their dependencies while reusing unaffected accepted artifacts.
S3 AWS object storage; proposed for internal source and artifact storage here.
GPU A processor commonly used to accelerate model inference.
Real-time factor Processing time divided by audio duration under defined hardware and workload conditions.
POP Period of performance: the contractual time window for performing the work.
SOW Statement of work: the document defining requested work and deliverables.
RFI Request for information: a market-research request, distinct from an award or accepted implementation.

For precise format definitions, use J2 or the 2002 standard according to the approved profile. General operational definitions in this table are explanatory additions.

Reference

Source library and reading guide

Open the original files and the primary references behind the wiki.

Download this article as MarkdownArticle permalink

Sources were checked for this local edition on October 1, 2026. Attachment citations link directly to the original SAM.gov downloads. Preserved local copies remain available as backups; live links may change. Nearby article citations identify the source of specific claims. Public references do not automatically supersede attachments or resolve procurement ambiguities.

Original opportunity attachments

Source Open the source What it contributes
Original user analysis Original analysis text Feasibility judgment, architecture proposal, model shortlist and five proof gates. Preserved verbatim; its old sandbox links remain historical text.
Draft SOW Word document · Local backup · Text extract Scope, functions, integration, support, ownership, documentation, phases and milestones. Use section numbers rather than fixed Word page numbers.
J1 1203:2018 construction specification · Local backup 31 pages; version 1.3, January 2019. Exact NLS construction rules.
J2 Z39.86-2005 R2012 · Local backup 95 pages; a different edition from the SOW's stated 2002 target. Its printed page numbers have a 12-page front-matter offset.
J3 QA201801 navigation instructions · Local backup Eight pages; version 2.0, December 2018. How book structure becomes navigation.
Opportunity SAM.gov notice 2027-NLS-0075 Original opportunity and attachment context. Check it for later amendments.

The attachment links above are public SAM.gov download endpoints, without session tokens or API credentials. If a direct download becomes unavailable, return to the opportunity attachment list or use a local backup.

The original analysis and SOW text extract are local supplementary files, not attachments hosted on SAM.gov. The SOW text extract is a convenience copy; layout and table relationships may be flattened, so use the Word original to resolve ambiguities.

PDF links with #page= use one-based file page numbers, not necessarily the printed number on the page. Browser support for opening at a specific PDF page varies, and SAM.gov may download a PDF rather than display it at the requested page.

Standards and production references

Primary reference Why follow it
NLS construction 1203:2025 Compare the newer public construction rules with attached J1.
NLS mastering 1202:2026 Master-audio parameters and heading preparation; clarify application to synthesis.
NLS delivery 1206:2025 Archive layout, ZIP details and Transfer Portal delivery.
NLS specification index Discover the complete applicable specification family. The indexed mastering date differs from the live page.
NLS 1205 protection PDF Catalogued November 2007 reference; direct retrieval failed during this review. Its clauses and governing revision were not independently verified here.
LOC AMR-WB+ format description Additional official format context and link to 1205.
ANSI/NISO 2002 The standard explicitly cited in the SOW and NLS construction profile.
ANSI/NISO 2005 R2012 Online counterpart to the attached later edition.
DAISY PDTB specification Concepts behind protected content and authorization/key material.

Models and reusable software

Reference Scope of evidence
Qwen3 TTS project and 1.7B Base model Documented variants, local loading and exact model-card license.
Chatterbox Official model options, controls and project license.
Kokoro 82M Model card and local usage; example emits 24 kHz audio, making mastering interpretation relevant.
Fish Audio S2 Pro Model capability and licensing pointer; commercial use terms need separate review.
Qwen3 ASR ASR and separate forced-aligner documentation; aligner limit is up to five minutes, not a claim about every ASR mode.
WhisperX Transcription/alignment implementation and documented limitations.
NeMo text processing Normalization and inverse-normalization foundation.
DAISY Pipeline conversion DTBook-to-DAISY3 options; its stated input is 2005-3 DTBook.

These sources document options to evaluate. They do not establish NLS acceptance or benchmark results for this project.

Historical integration evidence and AWS options

Hindenburg's NLS encoder support article and APH Book Wizard's revision history show previous integration with NLS components. They do not confirm current availability, supported operating systems, redistribution terms or cloud deployment.

AWS Batch GPU jobs and Step Functions with Batch document orchestration options. The Library's approved environment and actual workloads determine whether those options fit.

What changed from the original analysis

The wiki reorganizes the analysis into linked topics and adds explanation rather than treating every recommendation as a requirement. Source review added resource-audio timing exceptions, WAV-relative timing semantics, exact attached-edition limits, draft milestones, support/warranty terms, a missing-specifications list, the live mastering edition, and the distinction between S3 staging and the public Transfer Portal.

The original analysis remains available unchanged for comparison. No benchmark, cost quote, accepted architecture, contract interpretation or playback demonstration has been added as a claimed result.

Prototype planning references

The prototype adds an engineering plan, hardware allocations and test targets, all labeled as proposals. Primary sources below document model/runtime capabilities and available hardware, not NLS approval or measured performance. Checked October 1, 2026.

Reference Use in the prototype plan
Qwen CustomVoice Exact proposed preset-voice checkpoint and license label.
MLX-Audio Qwen adapter Community Apple Silicon runtime; conversion needs independent qualification.
Whisper turbo, whisper.cpp, faster-whisper ASR checkpoint and Mac/CUDA runtime paths.
Qwen forced aligner, Qwen 4B instruction model Optional alignment and constrained structural-assistance candidates.
Apple Mac Studio, MacBook Pro M5 Pro/Max Published memory/processor/storage configurations, not performance claims.
AWS G6, AWS G5 Published single-GPU instance specifications; approval and availability still need confirmation.
AWS Pricing Calculator Obtain a region-specific quote; no current price is assumed in the worksheet.