NLS Proposal WorkspaceRFI · Evidence · Delivery
Start here

NLS pursuit and proposal workspace

Prepare a credible RFI response and build evidence for a possible later competition.

Download this article as MarkdownArticle permalink

Purpose: build the strongest credible response now and the evidence needed to compete for a possible later contract. The immediate notice is RFI/Sources Sought 2027-NLS-0075, with responses due October 12, 2026 at 5 p.m. EDT / 4 p.m. CDT. It requests capabilities, rough pricing and SOW feedback; it will not itself result in an award. Verified notice and response instructions.

Proposed position: an inspectable production workflow that produces correct protected books, controls correction effort, covers both source routes and supports continued Library operation. These are themes to prove, not established advantages or a promise to win. The current evidence is a sourced plan and draft; company qualifications, prototype results and bidder-approved pricing remain missing.

Prepare the strongest credible response

Use the right evidence for each claim

Read current instructions first, then strategy, coverage, gaps, pricing and the draft. A response writer should cite verified capability and describe future work as proposed. A technical lead should use the pages below to produce proof. An executive deciding whether to bid should review rights, team capacity, protected-output feasibility and commercial exposure.

This Cloudflare site is public. The published version contains public research and a planning draft. Keep actual proprietary rates, private references, resumes and negotiation decisions in a controlled working copy.

Build the prototype

Start with the prototype implementation plan, then compare the local model stack, Mac/AWS hardware and sizing worksheet, and test corpus and acceptance gates. This is a proposed plan; no model benchmark, hardware purchase or prototype implementation is recorded here.

Explore the diagrams

From buyer need to a defensible claim: trace the requirement, proposed mechanism, evidence and bounded benefit before making a competitive claim.

Choose a reading path

Your goal Read these pages in order
Prepare the current response Opportunity → Strategy → Coverage → Draft
Evaluate the technical approach Book model → Ingestion → Speech models → Compiler
Plan a pilot Prototype plan → Models → Hardware → Acceptance tests
Estimate operational effort Quality review → AWS and operations → Open decisions
Check a claim Gaps and evidence → nearby source → Source library

Follow the production path

Two input paths feed one common book model. The clickable flow below opens the corresponding explanations.

How to interpret the claims

Source requirement means a statement found in the supplied draft SOW or a named specification. Edition and applicability still matter. Proposal means an engineering choice from the analysis or this wiki. Open decision means an answer is needed from NLS or from testing before the design can be fixed.

Public specifications are useful comparison material. A newer public edition does not by itself amend this opportunity. The standards page explains the 2002/2005 mismatch and the difference between the attached and current public construction specifications.

What to prove first

Produce a tiny, complete, protected book that passes the approved validators and plays correctly on the intended NLS systems. In parallel, test sustained narration on representative complete books and measure correction effort. These demonstrations address format interoperability and speech quality separately.

About this edition

Original technical edition prepared October 1, 2026; pursuit alignment updated October 2, 2026, from the supplied analysis, the four opportunity attachments, and linked primary sources. Added explanations include a glossary, explicit acceptance evidence, dependency tracking, milestones, and decision prompts. Nothing in this wiki records answers from NLS or actual performance results.

Wiki content works offline. Attachment citations open the original files on SAM.gov and require an internet connection. The source library also offers local backup copies. Editable articles are stored alongside it as Markdown. Keep the whole folder together when moving it. See the source library for provenance and the glossary for unfamiliar terms.

Prepare the response

Current opportunity and response instructions

Verified RFI stage, October 12 deadline, requested information and proposed workback.

Download this article as MarkdownArticle permalink

The immediate opportunity is an RFI, not a request for an awardable proposal. To compete for a possible later contract, submit a useful, credible market-research response now and build evidence that can survive later evaluation. The live SAM.gov notice 2027-NLS-0075, checked October 2, 2026, expressly states that no contract will be awarded from this RFI and that participation is voluntary, with no reimbursement of response expenses.

Verified response facts

Item Live notice as checked October 2
Title TTS for DTB Pilot - Request for Information
Notice type Sources Sought, updated
Agency and office Library of Congress, Contracts Services
Latest displayed update September 28, 2026, 3:34 p.m. EDT; attachments added
Response due October 12, 2026, 5 p.m. EDT / 4 p.m. CDT
Requested material Capability statement, rough order of magnitude pricing, draft-SOW feedback/questions and other useful solution information
Response route The notice directs material to both listed contracting contacts: James Robinson, jarobinson@loc.gov; Brenda Kinlay, bkinlay@loc.gov
Classification NAICS 541511; PSC DA01
Not established Final solicitation, scored evaluation factors/weights, definitive submission template/page limit, final price structure, confirmed bidders or incumbent

Use the live notice as the controlling source and recheck for amendments before sending. No page limit or mandatory file format was found in the displayed RFI; absence in this review is not permission to ignore later instructions. The blank set-aside field does not establish the structure of a future competition. Draft-SOW resumes due “with Proposal Submission” are future-proposal preparation items, not evidence that this RFI requests a formal proposal package.

What to prepare now

Prepare a concise capability-led response with a clear proposed solution, evidence of relevant delivery, a transparent ROM and prioritized SOW feedback. Use the proposed draft, and attach only verified supporting material. An optional short evidence appendix should make the response easier to assess; it should not bury the capability statement in this entire wiki.

Keep the current RFI response and future proposal as separate artifacts. When a solicitation arrives, extract every instruction and evaluation factor into a new compliance matrix, then restructure the offer to match. The current coverage matrix is a preparation aid, not a substitute for those future instructions.

Proposed response workback

These are internal target dates, not Government deadlines or scheduled automations. The remaining window is short; an unbuilt 16-week prototype cannot become a claimed pre-response accomplishment.

Target Work product Owner role
October 2–4 Confirm bidder/prime-or-partner position; collect qualifications, relevant references and actual capability evidence; cost assumptions Capture lead / business owner
October 5–6 Validate strongest technical claims, identify gaps needing a specialist partner, complete ROM and SOW questions Technical lead + estimator
October 7–8 Independent review of relevance, evidence, rights language, schedule and price logic Reviewer independent of the writer
October 9 Approve finished response, contact route and any proprietary markings; target early submission if authorized Authorized company representative
Before October 12, 5 p.m. EDT Recheck notice/amendments, send through the specified route and retain submission evidence Authorized sender

This wiki does not send a response, schedule reminders or contact the Library. A separate explicit instruction is needed to send anything.

Separate present capability from future work

Use three clear categories in every response: verified existing capability, proposed implementation, and dependency or unresolved decision. At present, this workspace establishes research, a technical plan, diagrams and a draft response. It contains no completed TTS prototype, NLS-player test, benchmark, bidder resume or customer reference. That distinction protects credibility and identifies exactly where to invest before a formal competition.

Prepare the response

Strategy for a competitive offer

Proposed buyer benefits, proof requirements, teaming and internal bid decisions.

Download this article as MarkdownArticle permalink

Our proposed case is lower delivery uncertainty with evidence the Library can inspect: correct protected books, manageable correction effort, both source routes, and a system the Library can operate and maintain under the required rights arrangement. These are proposed themes to prove, not established advantages over an identified bidder. The current RFI asks for capability, market information, ROM pricing and SOW feedback; it supplies no scored evaluation criteria. Opportunity status.

Make each claim earn its place

  1. Buyer requirementBooks that work in the Library’s playback environment.Source: draft SOW compatibility and quality obligations.
  2. Proposed mechanismControlled output profile, deterministic construction and independent validation.Describe exactly how the requirement would be met.
  3. Proof to produceProtected sample + profile/tool versions + actual player test record.Current status: planned; not demonstrated in this workspace.
  4. Benefit to assessLess uncertainty at acceptance and a repeatable release decision.Limit the claim to the profiles, devices and content actually tested.
A feature list alone cannot establish an advantage. Repeat this chain for correction effort, both source routes and independent Library operation.

A feature becomes useful in a response when it addresses a Library need, explains the proposed mechanism, produces evidence and states the practical benefit. Avoid “best voice,” “fully compliant,” “production ready,” “lowest cost” or “unique” unless the exact claim is supported. No probability-of-win score is justified by the evidence available here.

Proposed theme Buyer benefit to substantiate Evidence that would make it credible Current position
Correct delivered books Less acceptance and interoperability risk Tiny protected book, governing profile, independent validator results and target-player navigation record Planned; tools/player access unresolved
Controlled correction effort Predictable operator workload and production cost Full-book defect ledger, before/after repair demonstration and reviewer minutes per accepted hour Planned; no measurements
Both source routes with provenance Reuse commercial narration while preserving text fidelity Source-to-output mapping, commercial-audio example, text-origin example, audited announcement handling Designed; no prototype
Operational independence Continued use and maintenance after handover Rights inventory, clean installation, offline inference, source/configuration exports and an independent operator rebuild Proposed; rights and handover unproved
AWS integration and accountable delivery Fit with the existing environment and controlled operational risk Relevant past performance, approved interface plan, named staff, recovery tests and realistic phased estimate Bidder evidence not supplied

The strongest initial message is the combination and its proof. TTS, DAISY output, pronunciation controls and local processing already exist in the market; they are not sufficient novelty claims. Competitive reference offerings.

Turn the prototype into a bid investment

The 16-week prototype plan describes proposed execution after resources and dependencies are available. For the RFI window, prioritize evidence already in hand and a limited feasibility artifact only if it can be produced honestly and safely. Do not fund a broad polished UI before establishing the encoder/protector route and relevant delivery capability.

The next incremental investment should answer the hardest buyer question: can this team produce an inspectable protected DTB and explain what it costs to correct and operate? A short test book plus recorded limitations is more credible than a staged voice demo presented as complete compliance. A successful narrow artifact must still be described by its actual scope.

Team where the evidence is weakest

Evaluate a specialist accessible-book/DTB partner, an audio QA contributor familiar with long-form narration, and an AWS/security integration partner only where the prime lacks demonstrated capability. Choose partners for attributable past work, named available personnel, rights to deliver components, interface ownership and ability to participate in acceptance testing. A logo on a slide or an unconfirmed teaming conversation is not capacity.

The bidder must decide prime versus subcontractor positioning, control of integration and acceptance, workshare, responsibility for defects, and ownership/redistribution arrangements. Preserve the draft's non-subscription and ownership obligations in the proposed solution; a dependency that prevents continuing use after handover undermines the central offer.

Honest competitive discipline

Compare public capabilities and the evidence we can produce. Do not speculate about unnamed bidders' competence, claim an incumbent relationship without award evidence, or equate an open-source model choice with exclusivity. Do not derive competitors' pricing from our labor assumptions. Use the gap register to decide where teaming or proof closes a meaningful disadvantage.

Useful differentiation often lies in acceptance evidence and execution discipline: a common test corpus, exact model/profile provenance, accessible correction flow, bounded repair, source coverage and a reproducible handover. Recommend these measurable, supplier-neutral acceptance methods in SOW feedback because they help the Library compare solutions, not because they exclude another vendor's technology.

Bid and response decision gates

RFI participation gate: a response can proceed once bidder identity, factual capabilities and an authorized ROM are ready, even if the future system has not been built. Disclose proposed work clearly. Future prime-bid gate: confirm a qualified PM/Senior Engineer, required rights, a credible protected-output path, suitable past performance or teaming evidence, feasible schedule and supported price. If an essential capability remains unsupported, resolve it through a partner or reconsider prime positioning.

These are our internal business gates, not Government eligibility rules. The actual solicitation will determine mandatory qualifications, evaluation, representations and submission requirements.

Use this wiki as the evidence workspace

Every technical page now connects its subject to the response claim it supports and the evidence still needed. The coverage matrix maps draft requirements to response sections and proof. The draft provides buyer-facing language; the ROM makes the estimate inspectable. Keep actual proprietary rates, resumes, nonpublic customer details and private bid decisions in access-controlled working files. This published wiki contains public-source analysis and an unpopulated planning draft; it is not a private bid room.

Prepare the response

Competitive reference offerings and differentiation

Documented public capabilities and fair tests for proposed advantages.

Download this article as MarkdownArticle permalink

No bidder or incumbent has been confirmed for this opportunity. The offerings below are public references that could inform a competitor, a component choice, a baseline or a partner. Vendor descriptions establish documented capabilities, not independent quality results, contract suitability or intent to bid. An undocumented feature is unknown here, not proven absent.

What established offerings already do

Public reference Documented capability What we need to prove in our response
Hindenburg Narrator tutorials and feature comparison Explicit NLS DTB export, alongside other accessible formats and navigation/markup tools Do not claim others cannot make NLS books. Demonstrate the exact approved profile, protected output, AWS workflow, delivery rights and measured operator effort for our proposed system.
DAISY Pipeline TTS-assisted conversion, synchronized audio/text and pronunciation lexicons, using local and cloud engines Explain reuse versus custom work and verify NLS-specific protection/packaging. Local inference and lexicons alone are not novel.
DAISY Obi Open-source structured audio and accessible-publication production, editing, conversion and validation Benchmark reviewed commercial-audio navigation and operator correction. Source availability alone does not distinguish us.
Pratsam Producer Automated DAISY/EPUB speech-synthesis production, content comparison, manual review and cloud or dedicated service Existing market offerings already combine automation and QA. Demonstrate our own long-book fidelity, correction economics and Library handover terms.
Amazon Polly as an integrator's component Long-form voices, lexicons and speech marks; asynchronous synthesis writes results to S3 A voice service is not a complete DTB system. Compare accepted-output quality and lifecycle dependencies; do not presume open models are cheaper or that paid APIs meet the draft's rights terms.

Sources checked October 2, 2026. None establishes another provider's proposed price, staff, exact protected-NLS compatibility, contract rights or approved deployment for this acquisition. Confirm licensing and test access before evaluating any product.

Compete through inspectable evidence

Propose a common corpus and acceptance tasks that apply fairly to any supplier: produce protected samples from both routes; navigate on actual target players; account for source content; repair a seeded defect; measure review and compute effort; recover interrupted work; and install the delivered system independently. These tests turn claims into comparable results while leaving implementation choices open.

Comparison dimension What to record Avoid this misleading shortcut
Audio fidelity and listener suitability Whole-book defects, fatigue/clarity feedback, voice stability and final decoded audio A single polished short clip
Navigation and protection Same profile, metadata expectations, player matrix and actual protected artifacts A generic DAISY label or browser playback
Correction effort Reviewer minutes, repair time, fraction regenerated and missed defects on the same seeded issues An unsupported percentage saving
Cost All labor, tool, compute, review, failure/rework, support and warranty assumptions GPU cost compared with a vendor's complete service price
Independence and rights Actual continuing-use/redistribution terms and a clean handover test “Open weights” treated as ownership of all components
Integration and reliability Approved interface transactions, retries, restore and audit trail AWS hosting treated as complete integration or security authorization

Do not publish a scored vendor league table until comparable tests exist. Mark untested as untested and preserve failed runs. Future evaluation weights must come from the solicitation, not a favorable scorecard we invent.

Reuse and teaming decision

Before building each subsystem, evaluate whether an established component can satisfy the profile, runtime, rights and maintainability requirements. Reuse can reduce implementation work; integration and regression testing still cost effort. A qualified partner may close protected-DTB, accessible-reader testing or delivery-record gaps faster than another model experiment. No vendor named above is a committed partner.

Our proposed differentiator is the demonstrated contract-specific delivery chain and its correction/handover evidence. The strategy page connects that position to buyer outcomes; the gap register records why it is not yet a claim of superiority.

Prepare the response

Proposed RFI response draft

Editable response language with explicit placeholders and proposed commitments.

Download this article as MarkdownArticle permalink

Working draft for review — not submission ready. This is proposed response language for the current RFI and a starting point for a possible later technical proposal. It has not been sent. Bracketed fields require real bidder information. All implementation statements describe proposed work; no prototype, benchmark, NLS acceptance or corporate qualifications have been demonstrated in this workspace. Verify the current instructions and obtain company approval before submission.

An editable Word copy is available here: Download the proposed RFI response (.docx). This public version contains placeholders and an illustrative pricing scenario only; keep private bidder evidence in the controlled submission copy.

Response identification

To: James Robinson and Brenda Kinlay, Library of Congress Contracts Services
Subject: Response to RFI 2027-NLS-0075 — TTS for DTB Pilot
Respondent: [LEGAL COMPANY NAME]
Business address and authorized contact: [ADDRESS, NAME, TITLE, BUSINESS EMAIL, PHONE]
Relevant entity identifiers and business status: [VERIFIED INFORMATION AS APPLICABLE]
Proposed role: [PRIME / SUBCONTRACTOR / CONFIRMED TEAM AND RESPONSIBILITIES]
Response date: [DATE]

Executive response

[COMPANY] submits this capability and technical approach for the Library's market research on automated digital talking book production. We propose an inspectable workflow for both electronic text and existing commercial audio, designed around reliable navigation, faithful speech, protected output and continued Library operation.

Our proposed approach would establish the protected-output and target-player path early, then measure full-book fidelity and correction effort before scaling production. The Library would be able to trace source material to spoken content, inspect exceptions, authorize corrections and retain the evidence used to release each book. Model choice would remain replaceable behind stable interfaces.

This response describes a proposed implementation, relevant capabilities to be substantiated below, an illustrative ROM basis for review, and questions that materially affect scope and acceptance. We understand that the current notice is voluntary market research, is not a request for official proposals, provides no reimbursement of response costs and will not itself result in a contract award.

Bidder capability and relevant experience

[INSERT A FACTUAL COMPANY CAPABILITY PARAGRAPH: what the company currently delivers, its actual role in comparable software/audio/accessibility/AWS work, and evidence supporting each statement. Do not use the proposed architecture as proof of past delivery.]

Provide two or three relevant examples if available; this count is a drafting recommendation, not a stated RFI requirement. For each: [CUSTOMER/PROJECT AND DATES], [COMPANY OR INDIVIDUAL ROLE], [COMPARABLE TECHNICAL/OPERATIONAL CHALLENGE], [VERIFIABLE RESULT], [RELEVANCE TO THIS WORK] and [REFERENCE PERMISSION/CONTACT IN THE CONTROLLED COPY]. Attribute partner and individual experience accurately. If no comparable corporate project exists, disclose that and explain the verified personnel or teaming evidence available.

The proposed key personnel are [PM NAME] and [SENIOR ENGINEER NAME]. [INSERT ACCURATE QUALIFICATIONS AND AVAILABILITY.] We would align these roles with the draft's responsibilities and qualification requirements, including relevant degrees and experience, and provide required resumes with a future proposal. Additional audio QA, accessibility, security and integration support would be supplied through [VERIFIED STAFF OR CONFIRMED PARTNER ARRANGEMENT].

Proposed technical approach

For electronic text, the proposed system would ingest EPUB, Word, RTF and plain text, preserve original files and reading order, and maintain links from source spans to the approved spoken representation. Deterministic parsing would precede model-assisted suggestions. Uncertain headings, notes, tables or pronunciation would be exposed for review rather than silently rewritten or omitted.

For commercial audiobooks, the workflow would accept WAV, MP3, OPUS and FLAC, preserve supplied narration through approved processing, and use available catalog data, matching text and audio evidence to propose navigation. The system would synthesize only approved announcements and supplemental material. It would not invent unavailable page references or replace the original narrator without an authorized requirement.

Both routes would feed a common versioned book representation. Accepted text/audio and a frozen timeline would drive deterministic XML/metadata generation, approved encoding, PDTB/NLS protection and packaging. The output profile would record governing editions and decisions; outputs would be validated independently and tested on the agreed Library player matrix. Access to approved encoders, protection tools, keys, validators and reference material is an early dependency.

Our proposed first local narration candidate is Qwen3-TTS 1.7B CustomVoice, evaluated against Kokoro-82M with independent Whisper-based discrepancy checks and human listening. This is a candidate selection, not a quality or rights approval. The final engine and voices would be selected from measured fidelity, listener suitability, resource use, supportability and acceptable component rights. We would not assume that upsampling alone resolves synthetic-mastering requirements.

The prototype would use Apple Silicon for authorized development fixtures and a separately qualified Linux/NVIDIA worker in the approved AWS environment. A proposed initial sizing experiment uses 64 GB of Mac unified memory and one AWS L4-class GPU with 24 GB GPU memory. Actual capacity and operating cost would be determined by representative complete-book tests, including failed attempts and human review. These resource figures are unbenchmarked planning inputs.

Quality and evidence plan

We propose a small complete-book corpus covering both routes, all specified input formats through representative books or fixtures, two selectable approved voices and difficult held-out passages. The pilot would inspect the final encoded/decoded release audio, source coverage, metadata, navigation, mastering and protection. ASR and alignment would flag issues; neither would independently authorize release.

The demonstration would include a tiny protected reference book, a complete text-origin book, a commercial-audio example, a seeded pronunciation/navigation correction, dependent-artifact rebuilds and interrupted-job recovery. Measurements would report accepted audio-hours, defects, correction and reviewer time, compute/memory use and limitations. We propose involving experienced operators and blind or print-disabled readers in evaluating clarity, sustained listening and navigation tasks, with accessible participation and permissions arranged in the project plan.

We would provide the written certifications validating reports required by draft §3.3, with the responsible authorized signatory, exact standards/profile scope and supporting validation evidence identified. We would not certify untested compliance.

All claimed results would identify their corpus, software/profile/model versions, hardware, procedure and acceptance status. We would distinguish internal engineering gates from Government acceptance and retain unresolved defects or blocked tests in the report.

Delivery management and operational transition

The proposed plan would deliver the five specified architecture views and eight technical documents within 30 calendar days of POP start, followed by Phase I prototype/sample/demo/test artifacts within six calendar months. A 16-week engineering target is an internal planning allowance, subject to staffing and external dependencies; it does not replace the contractual deadline. Phase II integration/operational testing and Phase III deployment would follow the draft's prior-phase acceptance triggers and 90-calendar-day windows. Phase III would include operational ramp-up, production workflow management, performance monitoring and quality control at the agreed scale, as well as deployment and handover. Phase IV would cover ongoing production support, audio/metadata/output validation, security and configuration maintenance, backup/recovery, documentation, reporting and training through the remaining agreed POP. Production volume and human-review responsibility must be defined to resource these obligations.

The PM would own communications, dependencies and 30-day status/QC reporting; the Senior Engineer would own technical integration and evidence quality. Interface agreements, staffing and the complete output profile would be baselined early. The kickoff trigger should be clarified because the draft references both award and POP start.

Handover would include delivered source, build/deployment artifacts, approved model and voice assets where permitted, configuration/profile exports, manuals, training, test records and recovery/rollback procedures. An operator other than the original developer would demonstrate installation and the reference workflow. Support would include the required response within 24 hours and the specified 12-month warranty, with covered defect correction at no additional charge. Detailed support staffing and escalation would be part of the approved service plan.

Security and rights approach

The design would protect source and generated content through least-privilege access, controlled transfer/storage, separate protection-key management, limited content in logs, bounded egress, backups and verified retention/deletion procedures. Government source materials would be used only for authorized contract performance and disposed of according to the agreed requirements. Section H, approved AWS accounts/regions/services and integration boundaries are required to finalize the controls and cost basis.

We intend to meet the draft's requirement for Library ownership of the resulting TTS software, with no contractor-retained ownership interest and no required separate, recurring or follow-on software license or access purchase as a condition of continued operation. Third-party models, voices, parsers, encoders and protection tools would be inventoried and reviewed for compatible rights and the draft's non-infringement obligation. We request clarification of acceptable treatment of pre-existing/open-source components so that the final rights arrangement is explicit rather than inferred from a model license label.

Rough order of magnitude

Drafting note: the following is an illustrative scenario, not an approved estimate from [COMPANY]. Replace or expressly approve it before external use. At assumed blended selling rates of $150–$200 per hour, 5,200–7,200 hours across Phases I–III, a 15% labor allowance for uncertainty/warranty and $40,000–$110,000 in direct-cost allowances, the example totals $937,000–$1,766,000. The basis of estimate shows the arithmetic and assumptions.

The phase examples are $503,000–$878,000 for Phase I, $251,500–$490,000 for Phase II and $182,500–$398,000 for Phase III. An illustrative Phase IV unit is $12,000–$32,000 per month for 80–160 labor hours, plus approved incremental infrastructure/direct costs. The illustrative monthly hours are conditional on agreed production volume, review workload, monitoring and coverage; they do not establish capacity for the full ongoing production obligation. The Phase III allowance assumes a limited initial operational ramp and must be recalculated for the agreed production scale. Total support duration is not established. Covered warranty corrections would not be billed again as support.

[REPLACE WITH BIDDER-APPROVED ROM, DATE, CURRENCY, WORK BREAKDOWN, RATE BASIS, ASSUMPTIONS AND EXCLUSIONS.] The estimate should distinguish one-time delivery, operational services and infrastructure, account for Government-furnished resources, and state the effects of content scope, security requirements, tool rights, integration and acceptance. Any RFI figure is nonbinding market-research information, not an official quote.

Feedback on the draft SOW

We recommend clarifying the following to support comparable estimates and testable acceptance:

  1. Standards baseline: which editions and precedence govern the draft's 2002 reference, attached 2005 standard, attached construction specification and “current applicable versions”? Please provide the complete applicable specification set and explain “D standards.”
  2. Tooling and playback: which encoders, protectors, validators, reference DTBs, player devices/apps and key-management interfaces will the Library provide? Please identify supported operating systems, automation interfaces and deployment/redistribution permissions.
  3. Synthetic mastering and scope: how do initial-recording, heading-audio and mastering rules apply to generated speech? Which languages, multimedia profiles, legacy Word formats and complex content must be accepted?
  4. Security and integration: please provide Section H, approved AWS boundary/services, content-handling rules, retention/destruction requirements, storage/distribution interface contracts and acceptance responsibilities.
  5. Performance and quality: what content mix, volume, turnaround, listener criteria and allowable correction effort should inform sizing? We recommend a common corpus, protected-player tasks, source-fidelity checks and accepted-output/review metrics for supplier-neutral comparison.
  6. Rights and delivery terms: how should Library ownership coexist with approved pre-existing/open-source model and software components? Please clarify total POP, Phase IV scope, warranty/support separation and the kickoff trigger.
  7. Response and future evaluation: is there a preferred RFI format/length, and what future demonstration evidence would be most useful? We have not assumed unpublished evaluation factors or weights.

Suggested supporting material

The controlled response package should contain only verified evidence: company capability statement, relevant project examples, personnel summaries/resumes as appropriate, an approved ROM basis, and any actual test artifact with scope/limitations. Planned demonstrations must remain labeled planned. Do not attach the entire internal wiki as a substitute for a focused response.

Adaptation for a future proposal

If a solicitation is issued, replace this drafting structure with its actual instructions and evaluation order. Prepare an executive value statement supported by evidence, a complete requirement response, technical/quality/security volumes as required, management/staffing and transition plans, verified past performance, and approved pricing/representations. Reconcile every proposed pilot boundary with final scope and identify assumptions or exceptions explicitly. This RFI draft does not itself constitute a compliant future proposal.

Close the gaps

Draft requirement coverage and proposal evidence

Map requested information and draft obligations to response and proof.

Download this article as MarkdownArticle permalink

Use this matrix to turn the technical wiki into response material and later acceptance evidence. It maps the supplied draft, not an awarded contract or a final evaluation scheme. All implementation entries are currently proposed or unproved. A linked design page is not a passed test. Draft SOW.

Current RFI response coverage

Requested information Draft response location Completion test
Capability statement Bidder capability and relevant experience Real company identity and attributable examples; proposed capabilities separated from existing ones
Possible solution Proposed technical approach Both source routes, protected output, quality control and AWS integration addressed
ROM pricing Rough order of magnitude Bidder-approved effort/rate basis and uncertainty; clearly not an official quote
Draft SOW feedback Feedback on the draft SOW Prioritized, constructive questions that change acceptance, architecture or cost
Submission routing Current notice instructions Latest notice checked, authorized response sent to the specified contacts, record retained

Draft SOW coverage

Source Proposed response commitment Supporting wiki / required proof
§2 and §3.12 ownership / non-subscription Meet Library ownership of the resulting TTS software, no contractor-retained ownership interest, and no required separate, recurring or follow-on software license/access purchase Model rights, reviewed component/voice inventory, licenses and independent installation; legal/commercial review open
§3.1 commercial input Preserve supplied audio; derive reviewed navigation/metadata; generate approved supplemental speech; accept WAV/MP3/OPUS/FLAC Ingestion; actual sample and format fixtures still needed
§3.2 text input EPUB/Word/RTF/text processing, selectable voices, WAV/AMR-WB+, navigation and agreed multimedia outputs Prototype scope; format coverage and complete-book fidelity evidence planned
§§3.1–3.3 standards / protection / packaging Versioned approved profile; XML/OPF/NCX/SMIL, approved AMR-WB+ and PDTB/NLS1205 protection; written certifications validating reports under §3.3, with responsible signatory and evidence Compiler, standards; actual tool/player proof blocked on access and decisions
§§3.4–3.5 batch / scalability Durable jobs, resumability, controlled concurrency and transparent performance measurements Architecture, hardware; batch/recovery and accepted-hours benchmarks unmeasured
§§3.6 and 5.1 security / source handling Least privilege, protected source/output, controlled keys/egress, retention and disposal Operations; Section H absent, controls matrix and test evidence pending
§3.7 AWS / interfaces Integrate with approved NLS AWS/storage/distribution boundaries and retain standard export/import Operations; actual interface agreement, AWS access and integration trace required
§3.8 documentation / training / support Operator/admin manuals, training, patch process and support response within 24 hours Roadmap; staff roster, support process and walkthrough required
§3.9 warranty Correct covered defects without additional charge for the specified 12-month warranty period Pricing basis; reserve and responsibility needed; distinguish support fees from covered correction
§§3.10–3.11 reporting / QC Every-30-day status/QC reports; audio, structure, metadata, integrity and player checks before delivery Acceptance plan; templates plus completed evidence, not model self-certification
§3.13.1 / §4.1 design package Five specified diagram views and eight named documents within 30 calendar days of POP start Prototype deliverables; existing explanatory figures are not the complete accepted package
Phase I prototype description / §4.1 Prototype, DTB/voice samples, workflow demo and test report within six calendar months Prototype plan, test corpus; estimated engineering window is not Government acceptance
Phase II / III and §4.1 Integration/operational testing, then deployment and operation at agreed scale with workflow management, monitoring and QC; each within 90 calendar days after prior-phase acceptance Roadmap; full phase work and acceptance dependencies included in ROM
Phase IV Ongoing production/quality validation, security/configuration, backup/recovery, support, training and reporting through remaining POP Pricing; total duration unknown, monthly unit and scope needed
§5.2 / §4.1 kickoff Plan virtual kickoff promptly; clarify award-vs-POP trigger Decisions; retain ten-business-day requirement without inventing a start date
§6.1 key personnel Qualified PM and Senior Engineer, accurate resumes and availability Gap register; named personnel and qualifications not supplied

Turn coverage into a submission check

When the actual solicitation is issued, add its instruction identifier, volume/section, evaluation factor, page limit, responsible writer, evidence reference and review state. Do not invent evaluation weights from this SOW. Validate required representations, registrations, eligibility, amendments, due time and delivery method against that solicitation and the bidder's actual situation.

For each major claim, ask whether a reviewer can find the buyer need, proposed behavior, evidence and limitation in one place. Remove claims whose only support is another planning paragraph. Keep the engineering plan broad enough to satisfy final obligations; an English-only pilot or one multimedia profile is a proposed test boundary requiring agreement, not an exception already granted by the Library.

Close the gaps

Gaps to close before a credible response

Prioritized evidence, staffing, rights, cost and execution gaps with owner roles.

Download this article as MarkdownArticle permalink

The largest current gaps are bidder-specific evidence, protected-output proof and a supported cost estimate. The technical design is detailed, but no company qualifications, named staff, delivery references, tested prototype or approved commercial estimate have been supplied. Treat those absences as work items rather than fill them with persuasive adjectives. The RFI remains open as shown on the opportunity page.

Prioritized closure register

“Before response” means before an authorized RFI submission; it does not imply the completed production system is due then. “Before bid” means before committing under any later solicitation. Owners below are proposed roles, not assigned people.

Priority and gap Why it could cost us the opportunity Closure evidence and proposed owner Timing / present state
P0 — Bidder identity and position A technical concept without an accountable supplier is hard to assess Legal company name, contact, relevant registration facts as applicable, business status, prime/partner decision — business owner Before response; missing
P0 — Relevant delivery record The Library cannot assess execution credibility from architecture alone Two or three factual case studies: work scope, bidder's role, dates, comparable challenge, measurable result and authorized reference — capture lead Before response where available; not supplied
P0 — Named PM and Senior Engineer Draft §6.1 identifies key roles and qualifications Accurate resumes, availability, speech/audio/integration experience, degree and experience verification; partner commitment if needed — delivery lead Capability summary now; full required evidence before future proposal; missing
P0 — Credible ROM Unsupported price makes the market-research response less useful Actual staffing/rates/effort, direct costs, assumptions, uncertainty, warranty reserve and approval — estimator Before response; illustrative model only
P0 — Rights and continued operation A recurring software dependency or unclear deliverable rights could defeat a core requirement Component/voice/encoder inventory; rights assessment and proposed delivery terms — technical + procurement leads Risks disclosed now; resolved before commitment; open
P1 — Encoder and protector access Good WAV audio is not the required protected DTB Approved versions, OS/headless interface, licensing, key handling and known reference output — DTB engineer Ask now; prove before compliance claim; unavailable here
P1 — Playback evidence A demo might fail where patrons actually listen Target device/app matrix, test access, navigation record and final decoded-audio checks — QA/NLS counterpart Before qualified-output claim; missing
P1 — Complete-book speech and correction Short samples conceal omissions, drift and review cost Held-out passages, complete-book results, defect counts, repair/reviewer effort and voice approval — audio/ML lead Feasibility samples if actually available now; full prototype gate later; unmeasured
P1 — AWS access and security baseline Environment restrictions can change cost and architecture Section H, approved account/region/services, boundary/egress and interface decisions — platform/security lead Clarify now; obtain before deployment commitment; missing
P1 — Profile and content scope Standards ambiguity can invalidate outputs or underprice work Governing edition matrix, multimedia/language/content classes, synthetic-mastering interpretation — DTB lead + Library Prioritized RFI questions; unresolved
P1 — Commercial audio route A text-only offer leaves an explicit requirement unanswered Audio-only and text-assisted fixtures, provenance and reviewed navigation/announcement demonstration — backend/ML leads Include approach now; demonstrate in prototype; planned
P1 — Accessible operator workflow An inaccessible repair interface undermines usable automation Keyboard/screen-reader task tests with representative operators, documented fixes — QA/accessibility lead Plan now; test before acceptance; no UI built
P1 — Support and warranty capacity A prototype team may not cover accepted delivery obligations 24-hour response process, production/QC workload, on-call ownership, escalation, patch plan and separately budgeted warranty correction — delivery manager Price and describe before commitment; staffing unconfirmed
P2 — Recovery and independent handover Dependence on the original engineer weakens sustainability Interrupted-job/restore tests, clean installation by another operator, source/model/profile export — platform lead Prototype and later-phase gates; planned
P2 — Final solicitation instructions A strong narrative may still be nonresponsive New instruction/evaluation crosswalk, page/format limits, amendments and submission checklist — proposal manager Once issued; not available

What would make the offer defensible

For each proposed advantage, preserve a compact evidence package: claim, exact scope, date/version, responsible person, reproducible procedure, raw result and limitation. Distinguish bidder past performance from a partner's work and from a newly proposed capability. A reference must not suggest the bidder performed work it did not perform.

Add a named responsible signatory and evidence basis for draft §3.3 written certifications validating reports. The current plan contains no certification and no authority to certify untested output.

At least two independent review questions should be answered before release: Does every capability claim have the right evidence? and Does the proposed effort cover what we promise? A writer should not approve their own unverified compliance, pricing or past-performance statements.

Immediate clarification questions

Prioritize questions that materially change feasibility or ROM: governing specification editions; approved encoder/protector and available reference books; player access; synthetic mastering rules; Section H and AWS boundary; source-language/multimedia scope; volume/turnaround and acceptable correction effort; integration interfaces and responsibility; third-party component rights; total POP and support expectations. Ask whether the Library has a preferred response length/format before assuming one. The draft response consolidates buyer-facing wording.

Public and private working material

The public wiki can explain a proposed approach and source-grounded gaps. Store real resumes, customer contacts, pricing approvals, teaming terms and any marked proprietary content privately; replace public placeholders only in an authorized submission copy. The editable draft currently includes no invented qualifications. No claim in this register is a finding that the bidder is ineligible; company facts have not yet been supplied.

Close the gaps

ROM pricing and basis of estimate

Illustrative phase costs, support unit, assumptions and approval work still needed.

Download this article as MarkdownArticle permalink

The RFI explicitly asks for rough order of magnitude pricing for market research, not official quotes. A response should price the full proposed obligation and expose assumptions rather than equate a GPU bill with delivered software. Current notice.

The figures below are an illustrative planning scenario created for this draft. They are not bidder-approved rates, supplier quotes, market research on competitor prices or a price commitment. Actual rates, workshare, tools, security scope and support duration are missing. Replace or approve these inputs before using a numerical estimate in an external response.

Transparent illustrative scenario

Assume a blended fully burdened selling rate of $150–$200 per labor hour solely to demonstrate the calculation. This is not an assertion about the bidder's rate or the market. The rate includes labor burden, overhead and fee in this scenario. Add a separate 15% planning allowance on labor for uncertainty and warranty correction, then direct-cost allowances. Validate the allowance against actual risks; it is not a confidence interval.

Phase Assumed labor hours Labor range 15% labor allowance Assumed direct costs Illustrative phase total
I — design and prototype 2,800–3,600 $420,000–$720,000 $63,000–$108,000 $20,000–$50,000 $503,000–$878,000
II — integration and operational tests 1,400–2,000 $210,000–$400,000 $31,500–$60,000 $10,000–$30,000 $251,500–$490,000
III — deployment and operational ramp 1,000–1,600 $150,000–$320,000 $22,500–$48,000 $10,000–$30,000 $182,500–$398,000
I–III combined 5,200–7,200 $780,000–$1,440,000 $117,000–$216,000 $40,000–$110,000 $937,000–$1,766,000

Low scenarios use low hours/rates/allowances; high scenarios use high inputs. These are sensitivity cases, not a statistical range. Phase I hours encompass the proposed four-person engineering core, fractional PM/specialists and test participation; build a role-by-week staffing plan before asserting availability or the 16-week internal target. The six-month Phase I, 90-day Phase II and 90-day Phase III deadlines remain governed by their draft triggers and Government acceptance. Schedule.

Direct costs are placeholders for approved development/test infrastructure, test equipment/access, permitted third-party tools, accessible test participation and other project-specific purchases. Do not charge Government-furnished resources again. Record whether taxes, travel or procurement fees apply. No actual purchase, license approval or cloud capacity has been obtained through this worksheet.

Support and warranty need separate treatment

For Phase IV, an illustrative 80–160 labor hours per month at the same assumed rate yields $12,000–$32,000 per month, plus approved incremental infrastructure/direct costs. This is not priced 24/7 staffing or a resolution-time guarantee. The draft calls for response within 24 hours; validate coverage, escalation and workload before selecting a support model. Total Phase IV duration and total POP are unknown, so a responsible whole-contract total cannot be fixed from these attachments.

Phase III includes production workflow management, monitoring and QC at the agreed scale. Phase IV includes continuing production/audio/metadata validation, security, backup/recovery, reporting and training as well as incident support. The example assumes a limited initial operational ramp; the draft does not establish production volume or human-review workload. Recalculate Phase III and the illustrative 80–160 monthly hours against those inputs before claiming the full obligation is covered. These are not staffing guarantees for an unspecified production service.

Covered defects during the specified 12-month warranty must be corrected without additional charge to the Government. The internal risk/warranty allowance is a pricing provision, not a customer subscription or extra defect-correction fee. Separate maintenance/enhancement services from covered warranty work so the same effort is not charged twice. Draft SOW §§3.8–3.9, 3.12, §4.1.

Cost drivers to replace with evidence

A phase-based estimate must include PM/reporting, day-30 documents and diagrams, parser coverage, audio/TTS work, codec/protector integration, human QA, intended-reader participation, accessible UI, AWS interfaces, security controls, testing, training, deployment, rollback, warranty and support. Separate development-machine costs from production operation and Government-furnished services.

The most consequential sensitivities are review minutes per accepted hour, content/language complexity, legacy Word support, protection-tool automation/rights, security authorization, integration scope, benchmarked throughput and Government feedback/rework. A missing requirement should become an assumption and question, not vanish from the price.

The hardware worksheet estimates compute and review effort only. It is not this proposal ROM and must not be presented as total ownership cost. Treat the approximately 3,000-book annual background figure as context until duration, mix and turnaround are established.

Before issuing a ROM

Have an authorized estimator replace the illustrative rate and hours with a work breakdown, actual personnel availability, direct-cost basis and a risk register. Review third-party rights, ownership treatment and non-recurring access obligations. State the currency, estimate date, assumed scope, exclusions, range method, support unit and period covered. Clearly label the RFI estimate as nonbinding market-research information. Keep private margins, negotiated rates and internal risk tolerances outside this public wiki.

RFI preparation costs are borne by the responder under the notice and are not assumed reimbursable contract work. No competitor price or percentage saving has been inferred from this scenario.

Requirements and decisions

Scope and requirements

What the draft asks for, what it leaves open, and how to trace delivery evidence.

Download this article as MarkdownArticle permalink

The draft statement of work describes software for producing NLS digital talking books through two input routes. The engineering proposal is to share the downstream production system between those routes.

The draft background describes approximately 3,000 DTBs produced annually in the existing operation. That is context, not a numerical throughput acceptance target for the new system. SOW §1.

Response coverage

A competitive response needs an answer and an evidence plan for every material obligation, including personnel, ownership, reporting, support and warranty. Use the coverage matrix to map the draft to the response; use current RFI instructions for what is requested now. The technical scope below must not be mistaken for a published scoring scheme.

The two required input routes

Route Core work Important boundary
Existing commercial audiobook Ingest audio and metadata, identify divisions, create required supplemental speech, and construct the DTB. Proposed handling: preserve the supplied narration through approved mastering/conversion. A transcript supports analysis rather than replacing the recording.
Electronic text Ingest EPUB, Word, RTF, or plain text; recover structure; generate speech; construct the DTB. Preserve the source and document any change to its spoken form.

The commercial route explicitly accepts WAV, MP3, OPUS and FLAC, and requires AMR-WB+ conversion with MP3 encoding optional. The text route requires WAV and AMR-WB+ outputs, with MP3 encoding optional. These routes and formats come from SOW §§3.1–3.2. Scanned-book OCR is not an explicit core input in the draft. Treat it as a potential extension that needs scope agreement. A raw scan also creates recognition errors that the current text-first workflow would otherwise avoid.

Requirements and proposed evidence

The first two columns summarize the draft; the last is a proposed way to demonstrate completion, not a Government-approved acceptance checklist.

Area Draft requirement or concern Proposed evidence
Outputs NLS-compatible DTBs, navigation, metadata, required audio and related files; §§3.1–3.2 A reference book from each input route, validator results, and playback observations.
Processing Batch operation and automation with useful operator control Repeatable batch run, failed-job recovery, and an exception-review demonstration.
Integration Work within the Library's AWS environment and interfaces Deployment in the approved environment and a traced delivery to the agreed destination.
Security Federal security requirements referenced through Section H; §3.6 An agreed controls matrix after the missing section is supplied.
Compatibility and testing Library playback environment; §§3.11, 3.13 Test records for the agreed hardware, applications, and protected packages.
Ownership and continuity Non-subscription delivery and Library ownership; §§2, 3.12 Deliverable inventory, reproducible deployment, and dependency/voice-license records.
Documentation and support Technical documentation, training/support and warranty obligations Manuals, deployment runbook, handover exercise, support process and warranty terms.

Use the full SOW as the authority for exact wording; this table is a reading aid. The milestones page captures the concrete schedule terms.

Why local inference is the proposed default

The ownership and non-subscription language favors a system that can keep running without a required commercial speech API. The proposal is to deliver application code, approved model artifacts, configurations, and deployment machinery into the Library's environment.

That is an architectural interpretation, not a finding that every hosted service is forbidden or that open weights automatically satisfy the contract. AWS infrastructure still costs money. Third-party code, model weights, voice recordings, and commercial components each need their own rights assessment. Library ownership of delivered work does not erase third-party licenses.

Scope boundaries to settle

Agree which languages, book types, complex tables, equations, poetry, footnotes, and mixed-language passages are included in acceptance. Agree whether page navigation is required only when authoritative page references exist. Decide whether “various multimedia types” means an initial audio-plus-navigation product or additional full-text outputs. See open decisions.

A practical traceability record

For each accepted requirement, record its source section and edition, implementation component, test procedure, evidence artifact, and acceptance status. Keep this separate from model confidence scores. A model saying that a book looks correct is not proof of contractual acceptance.

Example proposal: the requirement for a valid navigation target maps to the compiler's reference checker, a deliberate broken-target test, and a playback test that reaches the expected chapter. This connects a rule to observable behavior.

Requirements that need explicit response language

Retain §3.3 written certifications validating reports, §6.1 qualified key personnel, and the complete ownership/non-subscription language in §§2 and 3.12, including no contractor-retained ownership interest or required separate/follow-on software access. Phase III includes operation at scale, and Phase IV includes production quality and maintenance activities. Use the expanded coverage matrix to prevent these obligations being compressed out of a technical narrative.

Requirements and decisions

Standards and output profiles

Resolve editions before treating DAISY or NLS compliance as one target.

Download this article as MarkdownArticle permalink

The supplied materials do not identify one unambiguous standards baseline. The draft repeatedly references Z39.86-2002, while the supplied ANSI/NISO attachment is the 2005 edition reaffirmed in 2012. Record and resolve this conflict before building an exporter.

Evidence of compatible delivery

The response value of a controlled profile is reduced acceptance uncertainty. Support that proposed benefit with the exact governing editions, resolved precedence, validator outputs and an actual protected-player record. Until those exist, describe a standards-control method rather than claim compliance. Track missing specifications and decisions in the gap register.

Editions in the source package

Source Identified edition How to use it
Draft SOW ANSI/NISO Z39.86-2002; applicable NLS specifications, with current-version language Starting statement of requested work; request an explicit precedence ruling.
J1 construction specification NLS 1203:2018, version 1.3, January 2019 Attached construction rules, including legacy limits and exact file behavior.
J2 ANSI/NISO standard Z39.86-2005, reaffirmed 2012 Useful DTB specification, but not interchangeable with the 2002 baseline.
J3 navigation instructions QA201801, December 2018 Rules for interpreting and marking book structure.
Public NLS construction page 1203:2025 at review Comparison reference; not an automatic replacement for J1.

The public 1203:2025 specification still references the 2002 standard and removes the older 250-file limit. The attached edition and the public edition therefore need separate profiles until NLS confirms which applies. Public construction specification.

Why the edition changes code

An edition can change element names, allowed structures, metadata, counts, or validation behavior. For example, the 2005 NCX has pageList and pageTarget; the 2002 standard uses navList and navTarget for page navigation. A generic DAISY exporter can produce valid output for its own target and still miss the required NLS target. Compare the 2002 standard with the attached 2005 standard.

The SOW's reference to multiple multimedia types also needs clarification against J1's audioNCX metadata requirement (§5.3). The proposal is to deliver the approved audio-and-navigation profile first while preserving internal text structure for possible future outputs.

Proposed versioned profile

Treat an output profile as configuration with an immutable version, not scattered constants throughout the code. It should identify:

  • Governing standards and exact schema/DTD/entity files, with hashes.
  • Required metadata, navigation classes, announcements, and supported multimedia types.
  • Master-audio and compressed-audio parameters, timing rules, and container constraints.
  • File names, file-size/count limits, checksum format, protection rules, and ZIP behavior.
  • Approved validator versions and target playback systems.
  • Documented NLS clarifications or authorized deviations, including their scope.

Every build records its profile ID. A change in profile creates a new build lineage so that later reviewers can explain why two otherwise similar books differ.

A profile selection process

  1. Send NLS a short matrix of conflicting requirements, with source sections.
  2. Ask for written confirmation of editions, precedence, output types, and supplied tools.
  3. Turn the approved rules into a machine-readable profile and human-readable checklist.
  4. Compile a small reference book and test it with the approved tools and players.
  5. Freeze that profile for the pilot; handle later amendments as explicit version changes.

These are proposed implementation steps. The wiki does not establish legal or contractual precedence between conflicting documents.

Current references are separate from the baseline

The public mastering, protection, and delivery pages should be checked alongside the approved profile. Preserve dated copies or hashes when implementation begins; a live link may change after this wiki is written.

Requirements and decisions

Open decisions and questions for NLS

A concrete register of answers needed before committing to implementation.

Download this article as MarkdownArticle permalink

The items below remain unresolved by the supplied materials. Suggested decision partners describe the kind of authority needed; they are not assignments made by NLS. Do not mark an item resolved without a written answer, accepted test result or other attributable evidence.

Separate buyer questions from bidder gaps

Use this page for technical clarifications to the Library. Use the gap register for internal company, staff, partner, pricing and proof work; use the opportunity page for current response instructions. Every material decision needs an owner role, requested-by date, cost/schedule impact and attributable resolution before it changes a proposal claim. Do not treat unpublished evaluation factors or an unconfirmed team as known.

Decisions that affect the architecture

Question Why it matters Evidence needed and suggested decision partner
Which editions and precedence govern the 2002/2005 conflict and current-version language? Changes XML structures, limits and test rules. Approved profile matrix from the contracting/technical authority.
Which multimedia types are required at acceptance? audioNCX and synchronized full text imply different outputs. Supported output-type list and reference books from NLS production/QA.
Which encoder, protector, validators and keys will NLS provide? Determines the feasible encoding/protection route and deployment platform. Current packages, versions, terms, headless interfaces, OS support and cloud approval.
What accompanies commercial recordings? Metadata and matching text determine recoverable structure and pages. Representative source sets and expected navigation from NLS production.
How do mastering rules apply to synthetic audio and generated labels? Native model rate, initial-recording language and copied heading rules may conflict with a chosen workflow. Approved interpretation and accepted synthetic example from NLS audio/QA.
What security controls and retention rules apply? Determines network, access, key, logging and deletion design. Missing Section H and approved architecture requirements from Library security.
What constitutes an accepted book? Controls quality thresholds, human review and pricing/effort assumptions. Content classes, language scope, player matrix, defect thresholds and throughput/turnaround targets.
What is the final submission interface for this project? The public Transfer Portal and proposed internal S3 staging are different interfaces. Integration contract, credentials/process ownership and acknowledgement semantics.
How are third-party model, voice and component rights handled? Library ownership/non-subscription terms must coexist with component rights. Approved inventory and applicable license/voice permissions from appropriate procurement/legal reviewers.

Specific mastering clarification

The live public mastering page is 1202:2026, revision 1.5, June 3, 2026. It specifies initial 44.1 kHz PCM and adds integrated loudness of −22 LUFS ±2; it also says heading clips are copied from primary audio. The public index still advertises a 2025 entry. Confirm the actual governing edition and how these rules apply to neural synthesis and separately generated navigation speech. Resampling a model's output alone does not settle the “initially recorded” interpretation. Current mastering page and specification index.

Additional source gaps

The SOW mentions 1201; J1 references 1208 for commercial-audio masters; only 1203 is among the four attachments. Request the complete applicable specification set. The phrase “and D standards” in SOW §3 is unexplained. The kickoff start trigger differs between §4.1 and §5.2. These are useful clarification items, not grounds to invent a preferred interpretation.

Proposed decision record

For each answer, store the exact question, governing source and section, answer date, responsible authority, evidence link, affected profile/version, affected tests and remaining conditions. Preserve superseded decisions as history.

Example: approving a new file-count rule updates the output profile and the corresponding validator. It should not quietly change an already submitted book or erase the rule used for its original build.

What can proceed while answers are pending

Build the source-preserving data model, importer experiments, independent quality checks and model evaluation harness. Prototype against a clearly labeled provisional profile. Avoid committing to a production codec/protection stack or promising numerical quality/throughput until the relevant components and acceptance rules are confirmed.

Build the evidence

Prototype implementation plan

A scoped build, proposed sequence, staffing, deliverables and decision gates.

Download this article as MarkdownArticle permalink

Proposed baseline: build a working, inspectable system that converts electronic text and supplied commercial audio into sample NLS digital talking books, with an operator able to correct a defect and rebuild the affected output. Start on an Apple Silicon Mac; prove the same workflow on one GPU in an approved AWS development environment. The proposed narration engine is Qwen3-TTS 1.7B CustomVoice, with Kokoro-82M as a compact comparison baseline; model selection remains conditional on listening tests and rights review.

Status: plan only. The wiki and its hosting are implemented; the production prototype, model benchmarks, purchased hardware, and NLS acceptance are not. Scope, staffing, durations, memory budgets and pass thresholds below are engineering proposals. Contract statements refer to the supplied draft SOW, not an awarded or complete contract.

Read model choices, Mac and AWS hardware, and test corpus and acceptance alongside this plan.

Use the prototype to substantiate the response

This prototype is the proposed proof investment behind the win strategy. Prioritize protected playback, source fidelity, measured repair effort and independent operation before expanding UI breadth. The current RFI deadline does not leave time to complete this 16-week plan; planned work must stay in future tense. The complete contract scope extends beyond this prototype, and pilot boundaries require agreement.

What the prototype must prove

The prototype should answer four questions: can the chosen voice read complete books faithfully; can the compiler produce a protected book that works on the intended players; can operators resolve exceptions without redoing a whole book; and can the Library operate the delivered system without a required paid speech API or recurring software-access license?

Phase I in the draft calls for an initial working TTS prototype, DTB samples, voice samples, an end-to-end workflow demonstration and a test report. The broader SOW covers both input routes. This proposed prototype deliberately exercises both so commercial-audio risks appear before Phase II. It does not imply that Phase I alone completes all integration, production or support obligations. Draft SOW §§3.1–3.2, 3.12–3.13; deliverable table §4.1.

Scope and boundaries

Area Included in this prototype Boundary or follow-on work
Text input EPUB, DOCX, RTF and UTF-8 plain text; preserve reading order, headings, source spans and metadata Confirm whether “Word Documents” includes legacy .doc; quarantine unsupported, malformed, encrypted or DRM-protected inputs rather than silently drop content. Scanned-page OCR is a separate extension.
Commercial audio WAV, MP3, OPUS and FLAC fixture coverage; preserve narration, propose navigation from available metadata/text/ASR, synthesize approved announcements Do not recreate the narrator or invent unavailable page references. Audio-only boundaries need review.
Languages and content English prose first, two approved selectable voices; ordinary headings, lists, notes, numbers and abbreviations English is a proposed pilot limit, not an SOW exemption. Mixed-language text, equations, poetry, complex tables and illustrations must be flagged and given an agreed reading treatment. Wider language support is a follow-on gate.
Output One agreed audio-plus-navigation profile, WAV masters, required XML/metadata, AMR-WB+, PDTB protection and delivery package MP3 is optional in the draft. Other multimedia types remain a scoped decision. A WAV/MP3 demo or unprotected package cannot substitute for required protected-output proof.
User workflow Import, inspect structure, select voice, queue job, inspect issues, edit pronunciations/labels, regenerate, compare versions, approve and export One production team and simple roles; no multi-tenant service, billing or public upload portal.
Automation Three-book batch, durable checkpoints, bounded retries, cancellation, resumable work and dependency-based rebuilds No promise to sustain the approximately 3,000-book annual background volume until representative benchmarks exist.
Deployment Native Mac model worker, shared API/data contracts, Linux NVIDIA worker in approved AWS test environment Real Library storage/distribution integration belongs to the agreed interface plan and Phase II acceptance. Prototype adapters may be simulated, visibly labeled.
Handover Source, pinned models and voices where redistribution is permitted, deployment recipe, test evidence and operator/admin instructions Production HA/DR, large-scale load qualification, full training and support remain later-phase work; backup/restore and access controls still receive prototype tests.

Build sequence and stop conditions

  1. 1
    Fix the test boundaryAgree profile, lawful fixtures, voices, tools and player access.
    ↓ supplies rules and reference evidence
  2. 2
    Make one tiny protected book playProve XML, timing, encoder, protector and real navigation together.
    ↓ establishes a working output path
  3. 3
    Complete both input routesFull narration from text; preserved commercial narration plus approved speech.
    ↓ provides complete books to inspect
  4. 4
    Measure, repair and recoverListen, flag defects, rebuild dependencies and survive interrupted jobs.
    ↓ produces qualified recipes and evidence
  5. 5
    Repeat in approved AWS and hand overDemonstrate reproducible operation; submit evidence for authorized acceptance.
Each gate earns the next claim. Import and model experiments may run in parallel; missing protection tooling or player access leaves the corresponding output gate blocked.

Planning allowance: 16 engineering weeks, assuming roughly four full-time engineering contributors plus fractional project management/specialists, available source rights, an NLS technical contact and timely access to required tooling and players. This is an internal estimate, not a delivery promise. Calendar commitments must follow the actual POP start and Government decisions. The draft sets a six-calendar-month Phase I deadline; do not equate that deadline with 24 weeks or assume acceptance/review latency is zero. Draft schedule.

Workstream and proposed window Concrete work and owner role Exit evidence / dependency
Weeks 1–2: baseline and risk probes Technical lead + audio/ML engineer: requirements matrix, provisional profile, corpus rights, two-voice short sample, protected tiny-book spike, Mac runtime probe Recorded decisions; encoder/protector/player access requested immediately. A missing codec/protection route blocks a claim of complete DTB success.
Weeks 1–4: design package Technical lead + security/QA contributors: five specified diagrams and eight specified documents Submit within 30 calendar days of POP start, regardless of the week labels. Unresolved decisions remain explicit.
Weeks 3–6: common core Backend/DTB engineer: source store, versioned book schema, EPUB/DOCX/RTF/text import, compiler rules and job state machine Tiny reference book round trip; no unaccounted source text or broken navigation; stable worker contracts.
Weeks 5–9: narration and review ML engineer + product/backend engineer: model adapters, pronunciation overrides, issue queue, transcript/audio comparison, versioned approvals Full text-origin book, both voices, deliberate pronunciation repair with correct downstream invalidation.
Weeks 7–11: commercial path Backend/DTB + ML engineers: media decoding, transcript alignment, metadata navigation, supplemental speech Complete commercial-source sample with matching text; audio-only sample with uncertainty surfaced for review.
Weeks 10–13: reliability and evaluation QA/accessibility lead + all engineers: complete corpus, malformed inputs, crash recovery, bounded retries, independent playback Measured defect and repair rates, all output checks, corpus coverage, accessible operator flow.
Weeks 13–16: AWS proof and handover Technical lead + platform/security contributor: approved AWS sandbox, controlled model transfer, integration rehearsal, clean rebuild and demo Same reference suite passes; compute/storage/review costs recorded; handover rehearsal and test report.
Remaining time before Phase I deadline Project lead coordinates Government feedback, component delays and corrective work Preserve contingency explicitly. If required access is late, revise the schedule and report the blocked gate rather than relabel a partial demo as accepted.

Some work overlaps; format/protection proof, approved output rules and lawful voice/model use constrain later release claims. Model experiments and source-preserving import work can continue while external answers are pending.

Concrete implementation shape

Use a modular application before distributing every component. Proposal: Python API and workers, PostgreSQL for state, an accessible browser operator interface, filesystem artifacts on Mac, and S3 artifacts in AWS. A small job queue and worker leases are sufficient initially; production orchestration can use approved AWS Batch/Step Functions after the basic contracts are stable. See architecture and operations.

Module Prototype behavior Durable contract / output
Importer Validate file type, sandbox document/media parsers, preserve originals, extract structure and source spans Source hash, parser version, ordered nodes, warnings, catalog claims with provenance
Book service Version source interpretation, spoken-text decisions, pronunciation rules, voice and output profile Immutable book version; every accepted passage mapped to a source span or an approved announcement
Job service Enqueue stage jobs, lease work, checkpoint, cancel and resume Input hash + recipe + stage identity; explicit failed/review/released states; no duplicate delivery
Speech worker Generate one bounded chunk under a book-level voice recipe PCM audio, duration, model/voice hash, runtime/precision, parameters, peak memory and elapsed time
Analysis worker ASR, alignment, silence/clipping checks and text-difference proposals Issues tied to exact text/audio versions; no autonomous acceptance
Compiler Assemble accepted masters; derive timing; build XML, encode, protect and package Profile version, actual delivered files, validator report, checksums and dependency graph
Review and delivery Human resolves exceptions; release requires current evidence; delivery adapter records receipt Actor, decision, rationale, approved version, package identifier and acknowledgement

Run local Metal/MLX inference as a native macOS process. Keep database/API containers optional. The Linux CUDA image is a separate build behind the same worker interface, not a claim that a Mac container image has GPU parity. Platform-dependent numerical differences are expected; compare coverage, quality and output validity rather than require identical generated waveforms.

Operator demonstration

Demonstrate a traceable session: import an EPUB, inspect recovered chapters and metadata, choose an approved voice, generate a chapter, inspect a flagged word, correct its pronunciation, audition before/after, rebuild the affected audio/timing/package, run validators, and navigate the protected result on a target player. Then repeat with a commercial audiobook, preserving its narration and adding only the approved generated material.

The UI needs book/job lists, a navigation-tree editor, text/audio issue review with keyboard controls, voice and lexicon settings, an evidence panel and a release action. A release blocked by stale validation should explain which change invalidated the evidence. Separate operator, reviewer/releaser and administrator permissions even if one person holds several roles during the pilot.

Deliverables and staffing

Propose four core contributors: a technical lead/platform engineer, a backend/DTB engineer, a speech/ML engineer and a QA/accessibility engineer. Product/UI design, audio editorial expertise, Library integration/security and procurement/license review are fractional support roles; they are real dependencies, not assumed free capacity. Reserve NLS reviewer and player access in the schedule. With one or two engineers, revise scope or time rather than inherit this estimate.

The day-30 design package should map directly to the SOW's five diagrams: System Content Diagram (wording in the draft), Component Architecture, Data Flow, DTB Packaging/Output Structure, and Security/Access. Its eight documents cover System Architecture, DTB Integration, Text Processing, Audio Output, Metadata/Packaging, Security/Data Protection, Testing/Validation, and Implementation Schedule. The wiki figures explain the proposal; they do not by themselves constitute all these signed-off deliverables. SOW §3.13.1 and §4.1.

Map the technical lead to the Senior Engineer key-personnel responsibility and allocate a separate Project Manager, proposed at 0.25–0.5 FTE during the prototype. The draft identifies both roles in §6.1; actual personnel must satisfy the stated qualifications. The PM owns the decision log, dependency escalation, schedule and status/quality-control reports every 30 calendar days. This management capacity is additional to the four engineering contributors, not hidden in their estimates. SOW §§3.10–3.11, 4.1, 6.1.

End-of-prototype delivery includes a reproducible code release, dependency/model/voice inventory and license notices, configuration/profile exports, corpus manifest, voice comparison, sample DTBs, test results, benchmark raw data, known limitations, backlog, setup/rollback/restore runbooks, a workflow recording and an operator/admin walkthrough. No source books, private recordings or keys go into the public wiki.

Decisions before implementation spending

Resolve the output-profile editions, synthetic-audio mastering interpretation, available encoder/protector and OS support, authorized voices, Library account/region/network access, Section H controls and retention, target players and acceptance owner. Obtain throughput and human-review budgets before selecting production capacity. Decision register.

If an encoder or protection tool is Windows-only, add a controlled Windows worker and a license/deployment review; Mac development does not remove that dependency. If full-book speech quality fails, compare the baseline and challenger or narrow the agreed content class; do not hide failures behind unbounded retries. If a model license cannot meet the ownership/access terms, replace the component before investing in its integration.

Build the evidence

Proposed local model stack

Qwen narration, comparison models and practical Mac-to-CUDA runtime choices.

Download this article as MarkdownArticle permalink

Start with Qwen3-TTS-12Hz-1.7B-CustomVoice for narration, and evaluate it against Kokoro-82M. This is a proposed starting point, not a finding that either model meets NLS requirements. Keep model inference local to the developer Mac or the approved AWS environment; the delivered workflow should not depend on a paid speech API. Training or fine-tuning a new model is outside the initial prototype scope.

Selection criteria that survive a model change

Screen candidates first for acceptable component/voice rights and reproducible operation. Then compare complete-book fidelity, listener suitability, correction effort, resource use and maintainability on the same corpus. Reject a model if a critical rights or fidelity issue cannot be resolved; retain a tested alternative rather than bind the offer to a fashionable name. The comparison plan makes clear that local TTS itself is not unique.

Primary narration choice

The exact upstream artifact is Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice. Its card lists Apache 2.0, preset voices, multilingual synthesis and style instructions. Choose CustomVoice, rather than Base, for the first prototype because it provides preset speakers without requiring a new voice-cloning workflow. The official examples use CUDA/BF16. Official Qwen implementation.

Start with restrained narration and short sentence-group chunks. Upstream lists Ryan and Aiden as English-native presets, both male. Include other approved English voices in the comparison if the pilot needs a broader selection; multilingual capability alone does not prove equal narration quality for every speaker/language pairing. Treat speaker choice as book-level configuration. Never silently switch engines or voices to repair one sentence.

The official checkpoint tree is approximately 4.52 GB at review time, including associated files. That is download size, not peak inference memory. The “1.7B” name alone cannot determine the total working footprint. Pin the selected revision and record actual memory during the runtime probe.

Supporting models and alternatives

Role Proposed artifact What it does, and its limit
Compact TTS comparison hexgrad/Kokoro-82M Small Apache-2.0 model with 24 kHz sample output. Useful for plumbing, announcements and voice comparisons; English voice catalog provides more preset choices. Its size does not establish full-book quality.
Independent speech check openai/whisper-large-v3-turbo MIT, 809M-parameter ASR candidate. Transcribes generated and commercial audio to flag discrepancies. Hallucinations and recognition errors mean its transcript is evidence for review, not truth.
Alignment experiment Qwen/Qwen3-ForcedAligner-0.6B Apache 2.0; aligns supplied text/audio, with a documented maximum of five minutes per input. Use much shorter overlapping windows and reconcile boundaries. Alignment cannot independently prove the supplied words were spoken.
Optional structure assistant Qwen/Qwen3-4B-Instruct-2507 Apache 2.0 instruction model for schema-constrained heading/reading-order suggestions on ambiguous input. Use a proposed 4–8K context initially. It must not rewrite narration, invent catalog facts or generate authoritative final XML.
Later challenger Chatterbox MIT project with several model variants and an embedded watermark. Compare only if the initial two-engine evaluation leaves a meaningful quality gap; validate exact variant, voice rights, runtime and watermark behavior after encoding.

The prototype needs TTS plus independent audio checking; it does not need every optional model loaded or even integrated on day one. Deterministic parsing remains the first choice for documents with usable structure.

Mac and AWS runtime paths

Stage Apple Silicon development Linux NVIDIA reference
Qwen narration MLX-Audio Qwen adapter, beginning with the documented BF16 CustomVoice conversion Upstream Qwen PyTorch/CUDA runtime; BF16 baseline with pinned dependencies
Kokoro comparison Upstream Kokoro using CPU or MPS as supported by the pinned version Same upstream model and code on CUDA/CPU
Whisper discrepancy checks whisper.cpp, Apple Silicon/Metal, with a pinned model conversion faster-whisper, with matching model source and recorded decode/precision settings
Optional aligner MLX-Audio conversion only after timestamp checks against known fixtures Upstream Qwen aligner as reference implementation
Optional structure LLM MLX-LM, evaluated 4-bit conversion Pinned local inference runtime and the same approved source model; benchmark separately

MLX-Audio is a community implementation. Its support is not a claim that the original CUDA package runs unchanged on Mac. Qualify the conversion on identical passages and on full-book continuity. Metal/MLX uses the Mac's shared memory; a CUDA worker has separate GPU and host memory. Hardware tiers.

For the first quality comparison, keep Qwen TTS at BF16. Quantized TTS is a separate, held-out experiment. Record runtime-specific output changes; do not assume a smaller checkpoint preserves fidelity. Use separate inference environments when dependencies conflict.

Guardrails in the adapter

Each synthesis request carries approved text, language, speaker ID, narration style, pronunciation policy, model and runtime versions, precision and a deterministic job identifier. Save audio duration, sample rate, text hash, elapsed time, memory peak and warnings with the output. Seeded generation improves traceability but does not guarantee identical audio across hardware.

The adapter must reject empty or incomplete output and detect truncation. Kokoro's pipeline implementation has bounded phoneme sequences and truncation-related behavior; enforce complete source coverage before calling inference, not merely by reading logs afterward. Use sentence boundaries and source spans to split material, and keep the distinction between synthesis chunks, navigation units and packaged files.

Pronunciation overrides change a separately recorded spoken representation. Do not feed arbitrary SSML to a model unless its adapter explicitly supports it. Check names, numbers, units and negation even when overall ASR word error rate looks good. Use the test plan for release evidence.

Rights, mastering and handover

Apache/MIT labels are useful starting evidence, not blanket clearance of model provenance, voice rights, dependencies, redistribution or the draft's ownership/non-infringement terms. Maintain a reviewed inventory of exact checkpoint revisions, conversions, voice assets, code licenses and any reference-speaker permission. Voice cloning, automatic character voices and style effects are deferred unless explicitly approved for the content.

Pre-fetch approved artifacts into a controlled cache, verify hashes and prove inference with network access disabled. Continued operation must not require a vendor account, hosted model endpoint or license renewal. The final rights arrangement still needs procurement review against SOW §§2, 3.9, 3.12.

Native synthesis sample rates and token rates are not output-compliance settings. Keep original generated audio, then apply the approved mastering/encoding profile. Upsampling to 44.1 kHz does not settle whether a model satisfies a requirement about initial recording; resolve the mastering question before making a compliance claim.

Build the evidence

Mac hardware and contract-aligned cloud options

Mac tiers, AWS worker sizes, cloud-permission boundaries and a sizing worksheet.

Download this article as MarkdownArticle permalink

Proposed starting configuration: an actively cooled Apple Silicon Mac with 64 GB unified memory and a 1 TB SSD, plus a short-lived AWS g6.2xlarge test worker with one NVIDIA L4, 24 GB GPU memory, 8 vCPUs and 32 GiB host RAM. Run model stages sequentially at first. This is a practical provisioning proposal for the model stack, not a measured minimum, throughput guarantee or purchase order. AWS G6 specifications.

Hardware supports the cost and delivery case

Use these configurations to obtain reproducible evidence and a transparent cost per accepted hour. Developer hardware is a separate expense from Government production capacity. The full ROM includes people, integration, testing, rights work, warranty and support; the GPU worksheet below does not price the whole contract. Do not claim lower total cost than another offer without comparable scope and results.

What the draft establishes about cloud providers

The supplied draft says to integrate with NLS AWS systems (§3.7), and describes existing S3-based storage (§1). It does not provide a cloud-provider allowlist or explicitly require every workload to execute exclusively in AWS. Because Section H is absent, the permitted account, region, services, data residency, network boundary and security authorization cannot be determined from these attachments alone. Draft SOW §§1, 3.6–3.7, 5.1.

Environment / provider Evidence of permission Proposed use
Library-approved AWS environment AWS integration is explicit; particular compute service/account/region approval is not established Primary candidate for cloud inference, storage and integration testing. Confirm service availability and GPU quota before booking work.
AWS GovCloud Not named in the supplied draft A conditional AWS option only if the Library specifies that boundary; do not assume it is required, sufficient or interchangeable with commercial AWS.
Azure, Google Cloud, Oracle Cloud No authorization established by the supplied attachments Do not include as approved runtime choices. Request a written determination before any contract-content processing or deployment there.
RunPod, Lambda, CoreWeave or other GPU hosts No authorization established by the supplied attachments Potential technical alternatives only after provider and data-handling approval; excluded from the baseline.
Developer Mac Proposed engineering environment, not an expressly approved place for Government content Public-domain, synthetic or explicitly authorized fixtures until content-handling permission is confirmed.
This Cloudflare Pages wiki Public documentation hosting already requested by the user Documentation only. Its deployment establishes no permission to host source books, protection keys, private recordings or inference there.

Cloud certifications alone would not establish contractual permission. Keep software-access rights and consumption-based infrastructure costs as separate questions; the draft's non-subscription language should be clarified for the final procurement arrangement.

Mac development tiers

Tier Proposed hardware Intended use and tradeoff
Reuse existing hardware Apple Silicon, 16 GB unified memory; adequate free disk API/UI/compiler work, Kokoro and short serialized smoke tests. Useful starting machine; not the buying target for sustained Qwen plus QA development.
Practical sequential-worker floor Apple Silicon Pro/Max class, 32–36 GB unified memory, 1 TB SSD One model loaded at a time, bounded chunks/context, few competing apps. Verify actual peak memory and swap before accepting it for whole-book runs.
Recommended desktop Mac Studio, M4 Max, 16-core CPU / 40-core GPU, 64 GB unified memory, 1 TB SSD Comfortable development headroom for speech experiments and application tooling. Apple lists this memory/storage combination; the recommendation is an engineering judgment, not an Apple TTS benchmark.
Portable alternative MacBook Pro, M5 Pro with 20-core GPU, 64 GB unified memory, 1 TB SSD; M5 Max optional Similar memory budget in a portable, actively cooled form. More GPU cores may help, but throughput must be measured on the selected configuration.
Larger workstation 96–128 GB unified memory Defer unless simultaneous model workers, a larger LLM or training becomes an evidenced need. Extra capacity is not required by this proposed scope.

Apple's published configurations: Mac Studio specifications and M5 Pro/Max MacBook Pro specifications. These establish supported configurations, not current inventory, delivery dates or prices. Confirm availability before purchase. An existing 48 GB machine is also a reasonable evaluation workstation; do not replace it before the memory/throughput probe.

The Mac's memory is shared by GPU, OS and applications; it is not equivalent to the same number of gigabytes of dedicated NVIDIA VRAM. Reserve at least 200 GB initially free for the pilot, plus a backup destination. Use native Metal/MLX workers; validate Linux/CUDA separately. Headphones for careful listening, representative NLS playback hardware and an accessible test browser/screen reader are as important to acceptance as the GPU.

Working-memory budget to validate

Process Proposed starting allocation Scheduling policy
Qwen 1.7B BF16 TTS 8–12 GB Batch 1, bounded sentence groups; measure cache/codec/framework peaks
Kokoro baseline 2–4 GB Load instead of Qwen for a comparison run
Whisper turbo ASR 3–6 GB Inspect accepted/generated chunks after synthesis; no required co-residency
Optional forced aligner 2–4 GB Short windows; reconcile boundaries
Optional 4B LLM at 4-bit, 4–8K context 4–8 GB Separate structural suggestion job; no whole-book context assumption
OS, database, API, browser and build tools Reserve roughly 10–16 GB on Mac Limit container memory and competing apps; watch memory pressure

All allocations above are estimates to test, not vendor minimum specifications. They depend on runtime, precision, input length, batches and conversion. Do not add every row as if all models must reside simultaneously. Conversely, do not use raw checkpoint size as total RAM. Reduce batch/context first, then serialize or unload workers before buying larger hardware. Model evidence and runtime choices.

AWS instances and surrounding resources

Candidate Published hardware Prototype role
g6.2xlarge — first choice to test 1 × L4, 24 GB GPU memory; 8 vCPUs, 32 GiB RAM One sequential inference worker; separate CPU compilation if host memory becomes tight
g6.4xlarge Same single L4 / 24 GB GPU memory; 16 vCPUs, 64 GiB RAM More host-side parsing/media/cache headroom; it does not double VRAM
g5.2xlarge — availability fallback 1 × A10G, 24 GB GPU memory; 8 vCPUs, 32 GiB RAM Benchmark against G6 if approved region/quota or quoted cost makes it preferable
g5.4xlarge Same single A10G / 24 GB GPU memory; 16 vCPUs, 64 GiB RAM Host-memory alternative; validate the exact backend independently

Hardware values come from AWS G6 and AWS G5 product tables; GB and GiB retain AWS's published units. Family support does not establish capacity, regional availability or Library approval. Avoid fractional-GPU instances for the first reproducibility baseline. No multi-GPU/H100 reservation is justified by current evidence.

Proposed supporting resources: one 4-vCPU/16-GiB CPU application/compile node (or equivalent approved service), PostgreSQL, a private artifact bucket, a private model cache, encrypted 250 GB scratch/EBS capacity for the GPU node, bounded log retention, and backup storage. The initial CPU/database services may share a development node; isolate them for later production requirements. Use durable storage for checkpoints; instance-local scratch is replaceable. Start with on-demand test windows and shut workers down when idle; test interruption recovery before considering Spot.

Same book contract, different execution environments

Package identical schema versions, text fixtures, voice recipes and output profiles for both platforms. Compare semantic fidelity, timing validity, memory and accepted throughput. Do not require byte-identical stochastic narration across Metal and CUDA. Test the approved encoder/protector on its actual supported OS; if it is Windows-only, budget a separate approved worker rather than assume Linux or macOS support.

Estimate compute and review effort

Change the assumptions to see the capacity effect. Defaults are examples, not observed model speed or an AWS quote. Each worker is a single-GPU instance.

Busy GPU processing32.80 GPU-h
Allocated instance time, total43.73 h
Compute wall time, approximate43.73 h
Human review effort, separate30.20 person-h
GPU instance subtotal onlyEnter a rate

No data leaves this page. Changing worker count changes estimated compute duration, not total work. Human review, CPU work, queues and service limits may prevent linear speedup.

Show the arithmetic and exclusions
Generated = text hours + commercial hours × supplemental %
Attempted generation = generated × generation multiplier
Busy GPU-hours = attempted generation × TTS factor + (commercial hours + attempted generation) × ASR/alignment factor
Allocated instance-hours = busy GPU-hours ÷ utilization fraction
Compute wall-hours = allocated instance-hours ÷ workers
Review person-hours = (text + commercial + supplemental hours) × review minutes ÷ 60
GPU subtotal = allocated instance-hours × entered instance rate

The generation multiplier approximates total replacement-audio work, not the maximum retry count. This excludes CPU compilation, optional LLM jobs, repeated full-source analysis, storage, network, engineering and review costs. It assumes enough independent jobs for the requested concurrency.

Measure before sizing production. Substitute observed per-stage factors, actual regeneration and review effort, and a current region-specific quote.

The worksheet uses explicit assumptions; its defaults are illustrative, not benchmarks. It models one GPU per instance and one selected GPU stage at a time. Generated hours include full text narration and supplemental commercial-audio speech. The generation multiplier covers additional synthesis and checking of replacement audio. Commercial original audio is checked once; repeated full-source analysis, optional LLM work, CPU encoding, queue delays and storage/transfer need additional budget. Effective utilization converts busy processing hours to allocated instance-hours; aggregate wall time divides by worker count, while total paid capacity remains unchanged in this simplified model.

Human review is separate and may dominate elapsed delivery. Pricing is deliberately entered by the reader because region, operating system, purchasing arrangement and date change it. Use a current AWS Pricing Calculator estimate for the approved region, and add CPU/DB, EBS/S3, backups, logs, network egress, player access, engineering and human-review costs. Set an account budget and worker time limit before any cloud experiment; no cloud instances are provisioned by this plan.

Storage sizing and buying decision

As an arithmetic example, 44.1 kHz, 24-bit mono PCM uses about 0.476 GB per hour: sample rate × bytes/sample × channels × seconds. Fifty hours with four retained PCM versions is about 95 GB before originals, models, intermediates and backups. This example is not an approved mastering profile. Stereo doubles that quantity; compression and retention change actual needs.

Buy or reserve only after a 30-minute representative narration/analysis run, a long-chapter run and the protected-output spike establish fit. Capture peak memory, sustained swap, per-stage RTF, reviewer effort and tool compatibility. Scale worker count only after a three-book batch and fault recovery pass. Approve production capacity against accepted audio-hours and required turnaround, not model parameter count alone.

Build the evidence

Prototype corpus and acceptance plan

Complete-book fixtures, measurable evidence, failure tests and the demonstration.

Download this article as MarkdownArticle permalink

These are proposed engineering gates, not Government-approved acceptance thresholds or completed test results. The prototype should produce evidence for both source routes, full-book speech fidelity, protected playback, repair behavior and operational handover. The draft requires quality checks and functioning Library playback; it does not supply numerical TTS quality or throughput targets. SOW §§3.5, 3.11, 3.13.

Build evidence a reviewer can trust

Associate each run with the response claim it tests, preserve raw failures and compare alternatives on the same material. Include a proposed small panel of blind or print-disabled readers and experienced operators, with accessible materials, permissions, compensation and time budgeted. Record listening fatigue, clarity at usual playback speeds and navigation-task success; no panel has been recruited or tested here. A planned test remains not tested in the evidence ledger.

Corpus with a reason for each item

Use public-domain, created-for-test or explicitly authorized content on developer machines. Record permissions, source hashes, language, duration/word count, expected navigation and known difficult passages. Government-provided materials require the approved handling environment. Proposed corpus: six complete books, about 30–50 source/output audio-hours in total, plus a short adversarial suite. The range is a collection target, not a statement about available books or a statistical guarantee.

Corpus component Coverage Human reference work
Four text-origin books One each in EPUB, DOCX, RTF and plain text; include prose fiction, nonfiction with notes, numeric/abbreviation-heavy material and weakly structured text Establish authoritative reading order, metadata, headings and approved spoken forms; distribute difficult features across books.
Two commercial-origin books One with matching text/metadata; one audio-only with limited metadata Reference chapter boundaries and announcements; explicitly mark unknown page information.
Format fixtures Short valid WAV, MP3, OPUS, FLAC plus all four text formats; corrupt, encrypted and misleading-extension examples Known success/rejection result for every fixture. Avoid lossy re-encoding merely to enlarge the corpus.
Difficult-passage set At least 100 approved excerpts with names, numbers, dates, abbreviations, notes, repeated words, punctuation, long sentences and language switches Reviewer-approved pronunciation and reading policy; hold out 20 excerpts from tuning.
Tiny protected reference book Three chapters, nested navigation, announcement, a note and one known page reference Exact expected metadata/navigation plus target-player checks; use an NLS-approved reference package where available.

Compare two approved voices and two engines on the same difficult-passage set. Select one recipe before the expensive complete-book run, retain the holdout set, and use a second reviewer to adjudicate disagreements. A six-book corpus demonstrates coverage; it cannot establish a universal error rate for all genres or languages.

Exit gates

Gate Proposed pass condition Saved evidence
Source fidelity Every source span accounted for; every omitted, expanded or normalized spoken passage has an explicit rule or reviewed decision Source-to-spoken-text map, coverage report, normalization diff; ordinary punctuation transformations recorded by rule
Book structure All required metadata and expected navigation targets are correct in the fixtures; no broken references or unsupported invented page numbers XML/schema and profile validation, expected-vs-actual navigation comparison
Speech quality Full human listening pass of final encoded/decoded release audio for the six-book corpus, including supplemental announcements; zero known unresolved omissions, hallucinated passages, repetitions or material meaning/pronunciation defects in release candidates Timestamped defect ledger, repairs and reviewer sign-off; report residual uncertainty and minor defects separately
Voice suitability Blind comparison of clarity, stability, fatigue and joins; NLS reviewers agree voice acceptance threshold before scoring Per-voice distributions, sample counts, reviewer disagreement and approval; do not claim a generic MOS score proves acceptance
Timing and mastering Measured outputs satisfy the approved profile's clip boundaries, levels, sample rate and encoding rules Automated measurements plus spot-checked waveform boundaries; timing explanation
Encoding and protection Actual AMR-WB+ plus PDTB/NLS protection using approved tools; output decodes/plays through the authorized path Tool versions, protected package hashes, validator results, authorized-key playback; never include keys in reports
Playback Every navigation entry in the tiny book and all top-level/representative nested entries in long books work on the agreed NLS hardware/application matrix Device/application versions, exact actions and observed results; browser audio preview alone is insufficient
Repair correctness Pronunciation, paragraph-length and heading-label changes rebuild every dependent artifact and preserve unaffected accepted chunks Dependency comparison, new timeline/checksums, old validation invalidated, post-rebuild playback
Reliability Three-book queue survives a worker kill, duplicate request, interrupted upload and restart without lost decisions or duplicate release Job-event trace, recovered artifacts, deliberate fault results; proposed retry cap: two regeneration attempts after the first
Operator accessibility Complete import/review/repair/release workflow by keyboard and agreed screen-reader combinations; focus, labels, issue announcements and audio controls usable Manual task test, assistive-technology versions and findings; formal standards baseline agreed with the Library
Security and handover Unauthorized release/content access denied; logs omit source content/secrets; unapproved egress denied; agreed retention/deletion policy verified across originals, intermediates, caches and backups; backup restores a reviewed book; clean installation reproduces the reference workflow Access, retention/deletion and restore tests, dependency inventory, offline model-loading test and another operator's setup log

“No known unresolved defect” means the agreed review found and resolved defects; it does not prove the model can never produce an error. Full listening is deliberately expensive for this small pilot so the team can measure what an efficient production review policy would miss. ASR agreement alone must never grant release.

Benchmark procedure

  1. Freeze model, voice, text normalization, chunking, precision, runtime, driver/OS and machine configuration. Record exact model revision hashes; do not compare mutable latest labels.
  2. Run cold-load and warm inference separately. Use fixed short passages first, then one complete chapter, then whole books. Record output duration and errors along with elapsed time.
  3. Measure TTS, ASR/alignment, CPU mastering/encoding/protection, queue wait, transfer and human review separately. Report p50/p95 chunk latency, peak memory, swap, failure and repair rates; include sample counts.
  4. Repeat representative runs, including long inputs and interrupted jobs. Count failed generations, discarded audio and repair attempts in cost, rather than measuring only accepted best runs.
  5. Repeat the same corpus on Mac and the chosen AWS instance. A faster machine is useful only if accepted output and repair effort remain acceptable. Qualify quantized and converted model artifacts separately.

RTF (real-time factor) = processing seconds ÷ resulting audio seconds for a stated stage. RTF 0.5 means that stage takes half as long as its audio; RTF 2 takes twice as long. Do not use time-to-first-audio or streaming latency as whole-book throughput. The hardware worksheet is a planning tool until measured inputs replace its examples.

Report accepted audio-hours per paid instance-hour, cost per accepted audio-hour, and reviewer minutes per accepted audio-hour. Suggested feasibility target for discussion: combined automated processing at or below the audio duration on the selected cloud worker, with no sustained memory pressure. This is a proposed test target, not a measured result or an NLS requirement; failure should trigger a sizing/quality tradeoff, not suppression of data.

Traceability and report format

Each requirement row should contain an ID, exact SOW/specification edition and section, implementation version, test fixture, expected result, observed result, evidence URI, defect ID and acceptance authority. Status values should distinguish not tested, passed, failed, blocked and not applicable with rationale. A missing player or protector is blocked, not passed.

The final report includes corpus coverage, raw measurements, voice comparisons, issue severity definitions, resolved and open defects, security/environment exceptions, playback evidence and the next-phase recommendation. Record both successes and excluded material. Acceptance should be attributed to the authorized reviewer; the engineering team's internal gates do not replace Government acceptance.

Demonstration agenda

Allow about 60–90 minutes as a planning assumption: inspect the corpus/profile and rights; create a short text-origin book; audition two voices; review and repair a flagged passage; inspect regenerated navigation and hashes; show the commercial-audio path; navigate protected outputs on target hardware; inject/recover one worker failure; and review benchmark and handover evidence. Pre-generated complete books complement the live short example so the demonstration does not mistake a polished excerpt for long-form validation.

Technical approach

System architecture

Separate AI interpretation from reproducible file construction.

Download this article as MarkdownArticle permalink

The proposed system has two front ends and one production core. Electronic text needs parsing and narration. Existing audiobooks need structure recovery and supplemental speech. Both converge on the same book representation, audio processing, compiler, validation, and delivery stages.

Architectural choices that support the offer

Use this architecture to explain accountability: one shared core covers both routes; source provenance supports fidelity review; bounded jobs support recovery; replaceable models support continued operation. Each benefit needs evidence, not just a diagram. The strategy and coverage matrix connect these mechanisms to required demonstrations. The five SOW architecture views still need a complete deliverable package, particularly security/access and packaging detail.

How the parts connect

Both input routes preserve their source evidence and feed the same book model, compiler and acceptance process.

Electronic text routeEPUB · Word · RTF · text Original text and source structure
parse and normalize
Source-linked book structure Preserve reading order, notes and page evidence.
Commercial audio routeAudio + available metadata Keep supplied narration through approved mastering.
transcribe, infer and align
Proposed navigation Use metadata, matching text and audio evidence.
converge on one versioned record
plan speech and retain existing audio
accept content and freeze the audio timeline
Deterministic constructionControlled by the approved output profile
XML + timing→Encoding→Protection→Packaging
inspect the actual package and playback
Independent acceptance Validators, target players and authorized review.
accept
Release and delivery Only accepted outputs reach the agreed NLS interface.
Proposed architecture. The input route changes preparation and speech generation; both routes use the same controlled construction and acceptance process. Failed checks return to the affected stage for repair.

Who does what

Work Proposed approach Reason
Recognize uncertain structure and ambiguous text Deterministic parsing first, model-assisted proposals when useful Source evidence stays available and uncertainty can be reviewed.
Generate narration and announcements Approved, replaceable speech models Voices and model families can be evaluated without rewriting the compiler.
Find possible spoken defects Independent ASR, alignment, acoustic checks, targeted review No single recognizer is a perfect judge.
Maintain IDs, source coverage, references and timing Typed data structures and conventional code These properties need exact, repeatable checks.
Construct XML, encode, protect and package Versioned deterministic compiler plus approved media tools Compatibility depends on exact rules and artifacts.
Accept a book Agreed validators, playback tests and authorized review Passing a model check is only one source of evidence.

The main modules

Ingestion stores the original files and parses format-specific structure. The book service manages immutable versions, source spans, metadata claims and approvals. Speech and analysis workers synthesize or inspect bounded segments. The compiler uses accepted content and a frozen timeline to emit the required files. Independent validators report defects. The review interface lets an operator resolve exceptions and trigger dependent rebuilds.

The delivery adapter publishes only accepted, complete outputs through the agreed NLS interface. It should not infer an interface from the fact that the environment uses AWS.

Separate three kinds of segmentation

A speech chunk is a unit that can be regenerated. A navigation element is a destination a reader can select. A delivered audio file is a packaging and playback unit. They do not need to have matching boundaries.

For illustration, one chapter could contain 80 synthesis chunks, 12 navigable sections, and one delivered chapter audio file. Those numbers are explanatory, not requirements. Keeping the concepts separate enables local repairs while respecting file limits and meaningful navigation.

Illustrative chapter · counts from the example below · not a timing scale

Three kinds of segmentation solve different problems. A correction can replace a small speech chunk while navigation and delivery remain organized at larger scales. The counts are illustrative, not NLS limits.

A book behaves like an incremental build

A pronunciation correction can affect several speech chunks. Regenerate those chunks, reassemble the affected masters, recalculate the dependent timeline, rebuild navigation/XML, encode and protect the changed content, then update checksums and packaging. Unaffected accepted synthesis can stay cached.

This is a dependency graph: each output records the exact inputs from which it was made. If an input changes, downstream evidence becomes stale. The build system should never keep a green validation result for an old version of a file.

Boundaries for AI assistance

Model outputs are proposals or generated media until verified. They must not silently delete source text, invent page numbers, substitute a narrator, write arbitrary final XML, or mark their own output accepted. Books and metadata are untrusted content, including instructions that appear inside them.

A useful model response is structured: a proposed chapter boundary, the source span supporting it, alternative interpretations, and confidence evidence. The compiler consumes an accepted decision, not a free-form conversation.

Small services before many services

The module boundaries can initially live in one application with separate worker processes. Splitting every module into a network service is not necessary to prove the approach. Establish stable data contracts, job IDs and reproducible artifacts first; deploy independently where isolation or scale justifies it.

Technical approach

Canonical book model and rebuilds

Keep source passages, spoken text, audio, navigation and evidence connected.

Download this article as MarkdownArticle permalink

The canonical book model is the system's shared record of what the book contains and how each part becomes audio and navigation. “Canonical” means that downstream stages refer to the same structured representation; it does not mean that an AI transcript replaces the original source.

Proposed invariant: every source passage is represented in the output or has an explicit, reviewable reason for exclusion. This is the strongest protection against quiet omission, duplication and paraphrasing.

Show why targeted repair matters

The proposed value is an auditable correction with less repeated work. Demonstrate a seeded omission, pronunciation change and heading edit: account for source spans, invalidate stale approval, rebuild all dependent timing/files and retain unaffected accepted audio. Compare repair and full-rebuild effort on the same case before claiming time or cost savings. This is evidence for controlled correction effort, not a benefit already measured.

Follow one passage through the system

Proposed data model · trace one passage through a build
Source passageOriginal text + stable location$1,205
→
Approved spoken text“one thousand two hundred five dollars”
→
Accepted audioVersioned asset + exact sample positions

First arrow: retain the normalization rule and approval. Second arrow: retain the model, voice and lexicon versions.

Source span + stable node IDPreserves hierarchy and reading order
→
Structure recordHeading, paragraph, note, table or page marker; parent + order
Structure + accepted audioResolves logical destinations into media
→
Two navigation referencesMain narration destination and separate spoken heading-label range
Approved text + observed audioASR is evidence of what a recognizer heard
→
Quality recordArtifact version, check, evidence, defect and resolution; ASR stays separate from approved text
Exact input versions + dependent checks
Auditable output buildProfile and tool versions, dependency records and package hashes identify exactly what was delivered.
One passage, connected records. This proposed model preserves the source while linking each transformation, media asset and approval to an exact version. The currency example illustrates normalization; it is not an NLS-prescribed phrase.

What to retain

Record Essential information Why it matters
Source artifact Original bytes, hash, format, ingestion version Reproduce and audit a conversion.
Source span File and stable structural location; original text or audio range Locate the evidence behind every transformation.
Structure node Heading, paragraph, note, list, table, page marker; parent and order Preserve reading order and navigation semantics.
Text transformation Original text, approved spoken form, rule, rationale, reviewer Distinguish faithful normalization from accidental rewriting.
Speech recipe Model/weights version, voice asset, lexicon version, settings Explain and selectively regenerate the audio.
Audio mapping Asset hash, sample rate, start/end sample positions, source span IDs Tie content and navigation to actual media.
Quality record Check type, artifact version, evidence, defect status, resolution Prevent stale approvals from surviving changes.
Output build Profile version, tool versions, dependencies, package hashes Explain exactly what was delivered.

This schema is a proposal. Its final field names and storage technology are implementation choices.

Preserve three forms of text

Store original written text, approved spoken text, and observed ASR text separately. Each answers a different question: what did the source say, what should be spoken, and what did a recognizer hear?

Example:

Original text:  $1,205
Spoken text:    one thousand two hundred five dollars
Transformation: currency/USD
Source location: chapter 4, paragraph 12, characters 18–24
ASR evidence:   populated only after examining the generated audio

A normalization rule should never convert “fifty” and “fifteen” into the same comparison token. Preserve high-consequence details such as negation, units, dates and proper names.

Use stable IDs and immutable versions

Give structural nodes stable identities across edits when they represent the same passage. Give source snapshots and generated assets immutable version IDs or content hashes. A node can point to a new approved audio asset without overwriting the old one.

A timestamp alone is an inadequate identity: earlier edits can shift every later timestamp. Store the logical destination and the resolved audio position, then derive the latter for each completed build.

Track two independent audio references

A navigation item points to a destination in the main narration and may also reference a spoken label in a separate headings file. Record both asset identities and timing ranges explicitly. Confusing the two timelines can produce a correct-sounding label that jumps to the wrong place.

Explore which work changes

Proposed dependency model · explore a change

A lexicon correction regenerates affected speech. Downstream timing and packages follow the changed audio; unrelated accepted synthesis can be reused.

Rebuild or rerunReuseDepends on scope
Original source + structureReuseRetain the source snapshot and stable passage identities.
Main narration chunksRebuild affectedRegenerate passages using the corrected lexicon entry; reuse unrelated accepted chunks.
Spoken heading labelsDepends on scopeRegenerate only labels that use the corrected entry.
Assembled audio + final timingRebuild affectedReassemble affected masters and calculate final sample positions.
Navigation + XMLRebuild dependentRebuild dependent SMIL/NCX references and durations from the new timeline.
Encode, protect + packageRebuild changedProcess changed media and update dependent protection, checksums and packages.
Validation evidenceRerun dependentValidate the changed dependencies and completed build; prior approvals remain tied to prior hashes.

Scope matters: a surname can have different pronunciations in different contexts. Record which occurrences the accepted correction covers.

A local edit can reach the final package. This simplified dependency order illustrates the proposed rebuild strategy. Narration and heading labels have separate assets and timing references. The approved profile determines whether a heading clip is copied or synthesized. The implementation must follow the actual dependency graph.

Worked repair example

  1. An operator corrects a surname in the book lexicon.
  2. Dependency lookup finds the passages whose spoken forms use that entry.
  3. Those passages receive new speech recipes and audio versions.
  4. The affected assembled masters and their final sample positions change.
  5. Dependent SMIL/NCX timing, durations, checksums, protection and packages are rebuilt.
  6. Validation runs against the new artifact hashes; the prior build remains auditable.

The system should record who accepted the correction, why it was made, and whether all occurrences should change. A name may intentionally have different pronunciations in different contexts.

Define cache identity carefully

Include source/spoken text, model and voice artifacts, normalization and lexicon versions, generation parameters, and relevant processing versions. Store the accepted audio itself. A random seed cannot guarantee byte-identical future synthesis across different hardware or software versions.

Technical approach

Source ingestion and navigation recovery

Preserve existing structure and make uncertain structure reviewable.

Download this article as MarkdownArticle permalink

Ingestion should retain source evidence before attempting to improve it. The workflow differs by input format, but every importer should produce source-linked content and proposed navigation for the common book model.

Prove both required source routes

A text-only demonstration leaves a material gap. Build an input-coverage record for EPUB, Word, RTF, text, WAV, MP3, OPUS and FLAC, with expected navigation and explicit uncertainty. Include a commercial recording with matching text and one without it. Measure the review burden when metadata is weak, and clarify legacy Word scope. Tie the result to the response coverage matrix and test corpus.

Electronic text formats

Format Recover first Watch for
EPUB Package metadata, reading order, XHTML structure, headings, links, notes, authoritative page mapping Visual styling that disguises semantic structure; broken reading order.
Word Heading styles, outline levels, lists, tables, footnotes and endnotes Direct formatting that looks like a heading but has no heading style.
RTF or legacy Word Convert in an isolated process while keeping the original and a conversion record Lost notes, tables, symbols, or outline information.
Plain text Delimiters and deterministic heuristics, followed by model-assisted classification Missing hierarchy and ambiguous headings; uncertain reading order.

The format list comes from the SOW §§3.1–3.2; the importer design is proposed. Do not flatten a structured source into plain text before extracting its relationships.

Main content determines navigation

J3 §3.1 directs navigation work to the main content rather than simply copying the table of contents. A table of contents may omit subheadings or describe them differently. Inspect the actual text and preserve evidence for every proposed division.

J3 also separates hierarchical navigation from nonhierarchical structures such as notes and segments. A note is not necessarily another chapter. The correct distinction affects playback behavior and must survive ingestion into the compiler.

A model can propose that a passage is a section heading, but the proposal should include its source location and supporting evidence. Printed page numbers should come from an authoritative page map or source, not model invention.

Commercial recordings

Use evidence in this proposed order: publisher metadata and chapter boundaries, matching electronic text when available, an ASR transcript, semantic classification, and precise audio alignment. Record conflicts rather than silently choosing whichever source arrived last.

For example, a model can distinguish a spoken chapter announcement from dialogue mentioning “chapter three.” The final destination still needs alignment and a waveform-level boundary check.

Keep the supplied narration through approved mastering and conversion. The ASR transcript is an analysis aid that can contain errors. Synthesize only the additional speech the approved output requires. Availability of a narrator's recording is not permission to clone the narrator's voice.

A recording cannot reliably reveal an unspoken print heading or an unannounced print page number. Mark these as unavailable, and agree how that input class will be accepted.

Written text to spoken text

Normalization expands abbreviations, dates, numbers and symbols while retaining original source spans. The sentence “Dr. Smith paid $1,205 on 03/04/2026” needs context: “Dr.” may require disambiguation, the currency needs expansion, and the date convention is ambiguous.

Use deterministic rules for well-defined cases and a reviewable proposal for contextual cases. A book-level pronunciation lexicon keeps repeated terms consistent. NVIDIA NeMo text processing is a candidate foundation for text normalization, not a guarantee that the whole book is correctly interpreted.

Model adapters and source safety

Each speech adapter should declare which controls it supports: SSML, phonemes, pronunciation dictionaries, language selection, reference audio, and style instructions. Unsupported controls must fail clearly or use an agreed conversion; they should not be silently ignored.

Treat model-specific control strings in source text as content. Escape or otherwise neutralize them before a model can interpret them as a voice or performance instruction. Retain the literal source and record the transformation.

What an importer should report

Proposed outputs include an ordered content tree, source coverage map, metadata conflicts, lost-format warnings, navigation proposals and unresolved passages. A successful file parse alone is insufficient: compare recovered content and reading order with the source before moving into speech production.

Technical approach

Speech generation and model evaluation

Compare candidates on full-book fidelity and repair effort.

Download this article as MarkdownArticle permalink

Use a replaceable text-to-speech (TTS) interface and evaluate several local models on the same material. The names below are candidates from the original analysis. None has been demonstrated here to satisfy NLS acceptance requirements.

Choose speech for listener outcomes

The model name is a replaceable component, not the central competitive claim. Compare final encoded audio over complete books for fidelity, consistency, fatigue and correction effort. Include experienced operators and blind or print-disabled listeners, with agreed tasks and ordinary listening speeds. Secure voice rights and preserve contrary results. The competitive references show why a short natural-sounding clip alone will not distinguish this offer.

Prototype model selection

The prototype model plan selects Qwen3-TTS 1.7B CustomVoice as the initial narration candidate, with Kokoro as a comparison baseline. It narrows the original broad shortlist into explicit Mac/CUDA runtimes, supporting models and evaluation conditions.

Candidate roles

Candidate Documented characteristics Proposed role
Qwen3 TTS Local inference and multiple model/voice modes; the named 1.7B CustomVoice weights list Apache 2.0 Proposed first adapter uses CustomVoice presets; Base reference-voice workflows are conditional. See the prototype model plan.
Chatterbox Official project lists an English Turbo model and multilingual options under an MIT project license Independent challenger for consistency and difficult passages.
Kokoro 82M Small open-weight TTS model; card lists Apache 2.0 Compact baseline for narration and announcements.
Fish Audio S2 Pro Model-card licensing requires attention, including separate commercial licensing Conditional evaluation candidate after terms are resolved.

License labels do not establish rights to every associated voice, dataset, dependency or deployment use. Pin the exact artifacts and record applicable terms in the handover inventory. This wiki's model descriptions are not a legal license clearance.

A book level narration plan

Before synthesis, select an approved voice, pronunciation lexicon, language handling, and restrained pacing/style settings. Preserve these across chunks. Faithful reading is the default: selectable voices do not imply automatic character acting, added laughs, paraphrasing or sound effects.

Test paragraph or sentence-group chunks. The original analysis suggests experimenting with 20–60 seconds of speech per chunk; this is a tuning range, not a requirement or measured optimum. Preserve sentence boundaries, and compare voice continuity, pronunciation, joins, and correction cost at different lengths.

Bounded generation and repair

  1. Synthesize a chunk from approved spoken text and a recorded recipe.
  2. Transcribe with an independent recognizer and compare meaningful content.
  3. Check alignment, clipping, silence, level and joins.
  4. Repair a specific discrepancy through a lexicon, text, boundary or generation change.
  5. Escalate after an agreed attempt limit rather than retry indefinitely.

Do not silently switch narrators for a difficult sentence. An engine fallback must preserve the approved voice or trigger a deliberate review and potentially a larger regeneration.

ASR and alignment candidates

Qwen3 ASR provides transcription and a separate forced aligner. WhisperX combines transcription with alignment and documents limitations. Long books need bounded jobs with reconciled overlap, rather than a single whole-book alignment call.

ASR asks what a recognizer hears. Forced alignment asks where expected text fits the audio. These are different checks, and neither establishes perfect narration alone. See the validation page.

A useful evaluation corpus

Include ordinary prose and long chapters alongside names, dates, quantities, negation, abbreviations, dialogue, notes, lists, tables, mixed-language passages, and repeated terms. Add passages from each supported input class. Keep tuning material separate from a held-out acceptance set.

Measure complete books as well as excerpts. Report unresolved defects per accepted audio hour, review minutes, repair attempts, throughput, and consistency across joins. Do not infer whole-book quality from a 30-second demo or time to first audio.

Voice governance

Keep approved voice assets, consent or usage rights, version, language scope and allowed transformations with the production recipe. A change to the voice reference can affect every generated passage and therefore should invalidate the appropriate cache and quality evidence.

Technical approach

DTB compilation and packaging

Turn accepted content into exact, validated navigation and media files.

Download this article as MarkdownArticle permalink

The compiler converts an accepted book model and finalized audio into the files required by an approved NLS output profile. It should serialize typed data and enforce explicit rules. Generating final XML directly with an LLM would make exact compliance harder to control.

The strongest early technical proof

A working protected book on the intended NLS player is stronger evidence than a generic DAISY export. Package the test input, exact profile, tool versions, output hashes, validator reports and player observations so the Library can inspect the result. Approved encoder/protector access and deployment rights are unresolved bid risks; historical vendor integrations do not establish our access, license or partnership.

What the main files do

File or format Role
OPF Describes metadata, the package contents and reading/playback organization.
NCX Defines navigable structures, labels and destinations.
SMIL Connects ordered playback structures to particular media clips.
WAV masters Accepted uncompressed source audio used as the timing reference.
AMR-WB+ in 3GP Compressed audio with NLS-specific codec and container constraints.
Resource files Supply approved alternative navigation-class labels when needed.
dtb.md5 XML checksum document for the DTB files, excluding itself.
DTD and entity files Local dependencies needed to interpret and validate the XML.

These descriptions summarize J1 §§3.1–3.10. Exact applicability depends on the selected profile.

Freeze audio before final timing

Proposal: track positions as integer samples, complete approved mastering and assembly, then freeze the accepted WAV timeline. Derive clip timestamps from that final timeline, not from rough ASR estimates or an earlier synthesis chunk.

In attached J1 §3.2.5, timestamps remain relative to the accepted WAV; codec offset compensation is assigned to the player. Do not silently shift XML to encoded-audio timing or compensate twice. J1 pp10–11.

Timing windows depend on clip type

Hatched band = permitted location of the clip boundary
Read from the speech edges outward. Each zero marks the adjacent speech boundary. Lead-in and tail use separate scales; the speech block has no time scale. Select one endpoint within each permitted band. These are boundary windows, not required amounts of inserted silence. If a valid boundary is unavailable, flag the case or approve an audio edit and rebuild timing.
Clip type Start before speech End after speech Attachment source
Main SMIL narration clips 80–120 ms 150–300 ms J1 §3.3.4, p11
Heading labels 80–120 ms 150–300 ms J1 §3.4.5, p15
Alternative-class resource audio 100–200 ms 150–500 ms J3 §3.3.3, p8

The first two rows also require the boundaries to lie in silence. Both clip endpoints must be present. A shared global “navigation padding” constant would miss the resource exception.

Use alignment plus local waveform analysis to identify the actual speech boundary. If the required margin is unavailable, flag the case or make an approved audio edit and rebuild dependent timing. Choosing the middle of a permitted window is a proposed default, not an NLS rule.

Main destinations and spoken labels

A navigation destination resolves into primary audio through SMIL. Its spoken label may refer to a different range in the headings audio. Keep both references explicit. Labels must represent the heading, not the whole section. J1 §§3.4.3.4–3.4.3.5, pp13–14; J3 §3.1.2, p6.

J3 requires at least two level-one navPoints: the first marks the opening title/author with class title/author, and the last marks the closing announcement with class close. J3 §3.1.1, p5.

Notes, note references, sidebars, pages, line numbers and segments need their own semantics. J3 requires segment navigation for entries in specified lists, and J1 says not to create navLists for segments. Treating every navigation object as a chapter loses these distinctions. J3 §§3.2.1–3.2.4, pp6–7; J1 §3.4.4, p14.

Exact rules from the attached construction edition

Rule in J1 Compiler consequence
No more than 250 files; §3.1.3 Count all required files under that edition. The newer public edition differs.
Split SMIL above 100 KiB; at most 50 SMIL files; §3.3.12 Measure actual UTF-8 serialized bytes; 100 KiB is 102,400 bytes. Follow the stated escalation rule if limits are exceeded.
At most 5,000 navPoints; contact monitor when exceeding 1,000; §3.4.3.6 Distinguish a hard limit from a consultation trigger.
Lowercase filenames and prescribed production IDs/suffixes; §3.1 Generate names from the profile and NLS-supplied identity.
XML diskcheck checksum structure; §3.9 Ordinary md5sum text is insufficient. Include each file except the checksum itself; do not list the checksum in the OPF manifest.
Include referenced DTDs and entities; UTF-8 XML; §3.10 Package dependencies and validate against an approved local catalog.
Duration follows SMIL playback, including specified custom-test content; §5.3(v) Derive duration from the compiled playback structure, within the stated ±1 second accuracy.

See J1 pp9–20. Do not apply the attached 250-file limit to a newer profile by accident. Permit only approved local XML dependencies; do not enable arbitrary network/entity resolution for book content.

Encoding is an early interoperability risk

AMR-WB+ is not ordinary AMR-WB. Attached J1 specifies mono constant bitrate, frame type 23 and ISF index 8, plus particular 3GP metadata and field constraints. A file extension alone proves none of those properties. J1 §3.2.3, p10.

Hindenburg's support article describes an NLS-supplied encoder; APH's revision history records NLS protection/validation integration. These are historical precedents. Obtain current components, terms, supported operating systems, automation interfaces and cloud-deployment approval before relying on them.

Protection is part of the book

PDTB protection involves protected content and associated authorization/key material. S3 storage encryption protects stored objects but does not create a protected talking book that an NLS player can authorize. See the DAISY protection specification and request the governing NLS 1205 profile and tooling.

Packaging and final submission

The public 1206:2025 delivery specification describes an outer ZIP with protected/unprotected DTB archives and WAV masters. It specifies uncompressed entries and local-header CRC/size values with ZIP bit 3 clear. A seekable staging file is a proposed way to control those headers, followed by archive inspection and extracted-file checks. Delivery specification §§3–3.2.

Its §4 names the NLS Transfer Portal. Internal S3 staging is a separate architectural choice; confirm the project-specific final interface. Keep QA logs outside the delivery package unless that interface expressly permits them.

Reuse existing DAISY components selectively

DAISY Pipeline's DTBook-to-DAISY3 script accepts a 2005-3 DTBook input and offers speech-related conversion options. It is a useful candidate for selected transformations. That documentation does not establish an NLS 2002, AMR-WB+, protected export path. Keep the NLS profile and acceptance tests explicit even when reusing its code.

Technical approach

Quality checks and operator review

Detect defects, repair bounded passages, and preserve acceptance evidence.

Download this article as MarkdownArticle permalink

Automation succeeds when defects are detectable and repairs are precise. A generated file is only an intermediate result until content, audio, navigation, format and playback checks agree with the accepted requirements.

The following layered QA design is proposed. The draft requires quality and compatibility testing but does not provide numerical narration-error or review-effort thresholds. Those must be agreed before the pilot is judged. SOW §§3.5, 3.11, Phase II.

Make quality evidence easy to assess

Present the quality-control plan as a repeatable release decision: define defect severity, assign independent review, preserve failed tests and show how a repair changes the evidence. Include intended-reader navigation and sustained-listening tasks alongside operator accessibility. The acceptance plan defines proposed tests; Government acceptance remains separate. No confidence score or planned test is a demonstrated competitive result.

From a candidate to an accepted release

Proposed control flow. The retry budget, acceptance criteria and review authority must be agreed. The draft calls for quality and compatibility testing; the numerical acceptance thresholds still need agreement. Draft SOW §§3.5, 3.11 and Phase II.

Six layers of correctness

Layer Example defect Proposed evidence
Source fidelity A paragraph is omitted or reordered Source-span coverage and ordering checks against the imported book.
Speech fidelity A number changes or a word is repeated Independent transcription, protected-token comparison and targeted listening.
Acoustic quality Clipping, an abrupt join, long silence or voice drift Signal checks plus excerpt and sustained-listening review.
Navigation A label announces one chapter and lands in another Reference resolution, alignment and interaction tests.
Format Broken XML references, wrong codec, bad checksum or ZIP headers Schema, semantic, media and package inspectors.
Playback Protection, skip, resume or speed behavior fails on a target player Recorded test matrix using actual agreed players/apps.

ASR evidence is not ground truth

An ASR mismatch can arise from a TTS mistake, a recognizer mistake, or a legitimate difference between written and spoken forms. Forced alignment fits expected text to an audio signal; it cannot independently prove every expected word was spoken.

Combine independent transcription, alignment, acoustic evidence and targeted checks for names, quantities, units, dates and negation. For high-consequence discrepancies, use another recognizer or a human listener. Agreement between models increases evidence but does not guarantee correctness. WhisperX's documented limitations illustrate missing timestamps and other alignment problems.

A review screen should show enough evidence

For each issue, present the source passage, approved spoken text, recognizer output, short playable excerpt, timeline/waveform, reason for the flag, affected build and repair history. Offer concrete actions: correct pronunciation, edit a proposed boundary, regenerate a passage, or accept a documented exception.

An exception should retain rationale and authorized reviewer identity. Acceptance of one passage should not automatically suppress the same error class throughout a book.

Defect states and bounded retries

A proposed workflow is detected → triaged → repair attempted → rechecked → accepted or escalated. Set an agreed retry budget per issue. Distinguish processing failures, which may be safely retried, from content ambiguities, which often need a decision.

Retest everything invalidated by a repair. Changing a chunk's duration can invalidate later clip times even when the later audio remains unchanged. Tie every check result to the exact artifact hashes and output profile.

What to measure

Metric Definition to agree Why it helps
Unresolved content defects Confirmed omissions, substitutions or repetitions per accepted audio hour Tracks fidelity more directly than general “naturalness.”
Navigation accuracy Correct labels and destinations divided by evaluated items, by type Prevents chapter success from masking note/page failures.
Review effort Active human review minutes per accepted audio hour Measures the practical cost of automation.
Repair burden Attempts and regenerated audio per accepted hour Exposes unstable models and costly chunk sizes.
Throughput Accepted output hours per elapsed production hour and per compute hour Separates machine speed from accepted-book delivery.
Operational reliability Failed/resumed jobs, incomplete packages and recovery time Tests production behavior rather than isolated demos.

These are proposed definitions, not measured results or contractual targets. Report the test corpus, input class, language, hardware, model version and accepted-output denominator with every result. Avoid hiding rejected books from quality statistics; report rejections separately.

Playback tests beyond the opening chapter

Exercise all navigation levels, notes/sidebars and skipping behavior, segment stepping, resume, transitions between audio files, and faster playback. Check both protected and unprotected artifacts where relevant. Record the player/app version and profile.

The draft expressly calls for testing with the Library playback environment. The exact device matrix and pass criteria remain open decisions. Accessibility of the operator interface should also be tested with keyboard and assistive technology as a proposed implementation practice.

Keep acceptance independent

Use known reference books, deliberate malformed examples, and a held-out evaluation set. An AI coding assistant can help generate tests and analyze regressions, but the acceptance criteria and reference outputs should be reviewed independently from the implementation that produces them.

Deliver and support

AWS deployment and production operations

A proposed deployment, realistic sizing method, and operational handover.

Download this article as MarkdownArticle permalink

The draft requires integration with the Library's AWS systems and standard import/export without proprietary lock-in. The components below are a proposed implementation, not a mandated stack. Confirm available services, network boundaries, security rules and delivery interfaces with the Library. SOW §3.7.

Demonstrate a system the Library can operate

The offer needs an accountable transition and service plan as well as AWS boxes. Identify interface owners, support coverage and escalation within the draft’s 24-hour response obligation, patch/change control, rollback/restore, training and independent installation. Distinguish no-charge covered warranty correction from separately priced maintenance. Include these responsibilities in the ROM, and qualify claims with actual team and integration evidence.

Prototype hardware and cloud boundary

See Mac hardware and cloud options for proposed development tiers, single-GPU AWS workers and an adjustable capacity worksheet. The supplied draft supports AWS integration; it does not provide an allowlist for other providers or settle account, region or service approval.

How deployment responsibilities connect

Operator / reviewerStarts jobs, resolves exceptions, authorizes release
uses the review application
Proposed components in the Library’s approved AWS environment
Application + control planeBook versions, jobs, approvals and output profiles
state
PostgreSQLStores state and decisions for the control plane
schedules jobs and records results
Workflow orchestrationStep Functions + AWS Batch, or an approved equivalent
dispatches bounded, resumable work
Model workersTTS · transcription · alignment
Pinned models and voices; GPU or suitable CPU
Compilation workersParsing · XML · media tools · validation · packaging
Isolated CPU jobs
workers read inputs and write versioned artifacts
S3 artifact storage
Originals→Intermediates→Staging→Accepted outputs
Delivery adapter
Submits only the accepted package; records acknowledgement
crosses the agreed delivery interface
NLS receiving workflowPublic 1206 guidance names the Transfer Portal.
The project-specific integration still needs confirmation.
Proposed deployment, not a prescribed AWS stack. The control plane manages decisions, workers process content, S3 stores artifacts, and a separate delivery adapter submits accepted books. S3 storage is not itself the confirmed submission interface. The diagram shows logical responsibilities, not a network-security design.

Proposed deployment components

Component Candidate implementation Responsibility
Control plane Typed application service Book versions, jobs, profiles, approvals and audit events.
Metadata and job state PostgreSQL Structured records, transactions and searchable status.
Artifact storage S3 Separate originals, intermediates, staging and accepted outputs.
Speech and analysis workers Containerized Python workers with pinned models GPU inference or appropriate CPU inference.
Compilation workers Isolated CPU containers Parsing, XML, media inspection and packaging.
Workflow orchestration Step Functions with AWS Batch, or approved equivalent Dependencies, bounded retries, checkpoints and review gates.
Operator application Accessible web interface Exception review, progress and authorized release.
Delivery adapter Agreed NLS integration Submit validated packages and record acknowledgement.

AWS documents GPU jobs in Batch and Step Functions integration with Batch. These capabilities support the design, but do not establish regional availability, account quotas, approval or required capacity.

Jobs should be repeatable and resumable

Use immutable inputs, explicit output versions and idempotent job keys. A retried job should not duplicate a submission or overwrite an accepted book. Record checkpoints after durable artifacts are written and verified. Use separate states for failed processing, unresolved content, ready for review, and ready for delivery.

Keep model workers warm over a batch where appropriate. Loading weights for every paragraph wastes time, but the right worker lifetime depends on measured throughput and utilization.

Capacity without invented benchmarks

A useful planning equation is:

GPU-hours of work = source audio-hours × processing factor × retry multiplier
Wall-clock hours ≈ GPU-hours of work ÷ (GPU count × effective utilization)

Here the processing factor must be measured in GPU-hours per source audio-hour for the chosen workload. If it is a combined factor, avoid counting transcription or repair twice.

For illustration only: 30,000 hours × 0.5 GPU-hours per hour = 15,000 GPU-hours before retries. At a 1.2 retry multiplier, this becomes 18,000. Ten GPUs at 75% effective utilization would imply about 2,400 elapsed hours of compute processing. These are assumptions and arithmetic, not observed performance or a calendar delivery forecast. Human review, CPU work, queues and delivery can remain bottlenecks.

Benchmark the two input paths separately. Commercial recordings may need full-book transcription/alignment but only supplemental TTS; electronic text needs full narration. Measure accepted output, including failed attempts, rather than quoting isolated generation speed.

Security and source handling

The draft references federal controls in Section H, which is absent from the supplied attachments. It also calls for source/output protection, retention/destruction procedures and security documentation. Obtain that missing material before fixing the controls baseline. SOW §§3.6, 3.13.1, 5.1.

Proposed defaults include isolated inference in the approved environment, restricted unnecessary egress, pinned dependencies, least-privilege roles, separate content-protection keys, and limited source text in logs. Book content must be treated as data, including text that looks like instructions. SOW §5.1 limits use of Government source materials to contract performance and requires secure disposal at contract end. Implement those obligations alongside the agreed retention and deletion procedures.

Storage encryption, transfer encryption and PDTB protection solve different problems. Keep their responsibilities and key handling explicit.

Handover and maintenance

Deliver source code, deployment instructions, configuration, dependency/model/voice inventory, operator and administrator guides, recovery procedures, validation tools and a known reference build. Confirm which third-party components can be redistributed and which must be supplied by NLS.

The draft specifies support response within 24 hours and a 12-month defect-correction warranty after acceptance; it permits ongoing support/maintenance separately from prohibited recurring access requirements. See the milestone page and SOW §§3.8–3.9, 3.12.

Release evidence

For each released software/profile version, preserve migration notes, a reproducible deployment recipe, dependency hashes, regression results and the reference-book outputs. That makes later maintenance possible without relying on the original contractor's environment.

Production obligations to resource

Phase III extends beyond installing the system: the draft calls for production operation at scale, workflow management, monitoring and quality control. Phase IV includes ongoing production/quality validation, security and configuration management, backup/recovery, reporting, documentation and training. Define volumes and operator responsibility before claiming a small support retainer covers the work. The ROM exposes those assumptions, and the coverage matrix retains the written-certification obligation.

Deliver and support

Proof gates and draft milestones

Separate the proposed technical sequence from the actual SOW schedule.

Download this article as MarkdownArticle permalink

Use technical proof gates to reduce uncertainty while meeting the draft's separate contractual phases. The five gates below come from the engineering proposal; they are not the four phases named in the SOW.

Two schedules serve different purposes

The RFI workback governs preparation for the current October 12 response. The schedule below concerns proposed post-award delivery and the draft’s relative milestones. The 16-week engineering plan is not a pre-response commitment or the entire contract duration. Price the day-30 documents, acceptance dependencies, later phases and support in the basis of estimate.

Detailed prototype work plan

The prototype implementation plan adds a proposed 16-week engineering sequence, scope boundaries, owner roles and handover artifacts. The acceptance plan defines the proposed corpus and evidence. These planning estimates do not replace the draft deadlines below.

Proposed proof gates

Gate Work to demonstrate Pass evidence
1. Format and protection Build a tiny complete book with the approved profile, encoder, protection and packaging Approved validation results and working navigation/playback on target systems.
2. Long-form speech Compare candidate engines on chapters and complete representative books Agreed fidelity and quality results, with measured correction effort.
3. Both ingestion paths Process structured text, less-structured text, commercial audio with and without matching text Source coverage and independently checked navigation, including explicit unavailable information.
4. Repair and rebuilds Change pronunciation, replace a paragraph, edit a heading, interrupt processing and fail encoding Only necessary work repeats; all dependent timing, hashes, metadata and packaging remain correct.
5. Production operation Run batches, security checks, recovery and real integration Repeatable accepted outputs, measured throughput and operational handover.

Prove codec/protection compatibility early with simple known content so that format errors can be separated from speech errors. Speech and importer experiments can proceed in parallel after their required inputs are available.

Schedule in the supplied draft

POP means period of performance. No actual start date or total duration is supplied here, so the dates below are relative triggers rather than calendar commitments. SOW §4.1 and named phase descriptions.

Deliverable or event Draft timing
Key personnel resumes At proposal submission.
Kickoff Within 10 business days; the table uses POP start while §5.2 uses contract award. Clarify the trigger.
Five architecture diagrams and eight technical documents Within 30 calendar days after POP start.
Status and quality-control reports Every 30 calendar days during POP.
Phase I prototype, sample DTBs, workflow demonstration and test report Within six months of POP start.
Phase II integration and operational testing Starts after Phase I completion and Government acceptance; completed within 90 calendar days.
Phase III full deployment Starts after Phase II completion and Government acceptance; completed within 90 calendar days.
Phase IV ongoing support After Phase III completion and acceptance, through the remainder of POP.
Warranty 12 months after written Government acceptance of all deliverables.

The 24-hour support-response requirement appears in §3.8. The warranty includes free defect correction and patches under §3.9. Do not turn the phase durations into a promised total project length: Government acceptance timing and the total POP are not established here.

How the gates can fit the phases

Proposed mapping: use Phase I to resolve the profile and prove format/protection, model behavior and initial ingestion. Continue hardening ingestion, repair and operational behavior through Phase II integration tests. Phase III adds deployment, handover and acceptance evidence. Phase IV covers production support and controlled improvement.

This mapping does not reduce the deliverables required in any phase. Validate it against the final solicitation and approved project plan.

Early documents should expose real decisions

The first-month architecture package should show both source paths, data ownership, security boundaries, failure/retry handling, and the split between AI proposals and accepted compiler inputs. List unresolved standards and supplied-component dependencies rather than hiding them behind a generic cloud diagram.

Suggested readiness check before a pilot

Confirm the output profile; acquire reference books, target players and required tooling; agree supported content classes and quantitative acceptance criteria; record authorized voices and dependency rights; obtain Section H; and identify who can accept exceptions and release a book. These are suggested prerequisites tied to the decision register.

Reference

Talking book glossary

Plain-language definitions for the terms used throughout the wiki.

Download this article as MarkdownArticle permalink

This glossary explains the terms as used in this proposed production system. The linked articles and standards contain the implementation detail.

Term Meaning
NLS The Library of Congress National Library Service for the Blind and Print Disabled.
DTB Digital talking book: audio and supporting structure that allow accessible reading and navigation.
DAISY Standards and technology for accessible publications, including structured talking books.
ANSI/NISO Z39.86 The talking-book standard cited in the procurement; the edition matters.
audioNCX A multimedia type describing full audio with navigation structure; it does not by itself promise a synchronized full-text product.
TTS Text to speech: generating spoken audio from text.
ASR Automatic speech recognition: generating a transcript from audio.
Forced alignment Estimating where a supplied transcript fits the audio; distinct from independently recognizing what was spoken.
Canonical book model The shared structured record linking original sources, approved text, navigation, audio and evidence.
Provenance Records of where an item came from and how it was transformed.
Source span An identifiable passage or audio range in an original source.
Spoken form The approved text representation intended for narration, such as expanding a currency amount.
Lexicon A book-level or shared set of pronunciation decisions.
Inference Running a trained model to generate or analyze content.
Model weights Numerical parameters of a trained model; they are distinct from application code and voice assets.
Deterministic compiler Conventional software that produces prescribed files from explicit accepted inputs and a versioned rule set.
Output profile The exact standards, media, metadata and packaging rules chosen for a build.
OPF The package document describing metadata, included resources and organization.
NCX Navigation Control file for XML; defines navigation structures and labels.
SMIL Synchronized Multimedia Integration Language; links playback structures to media segments.
DTD Document Type Definition; a schema mechanism used by these XML formats.
Entity An XML declaration/reference whose allowed definitions must be controlled and packaged as required.
navPoint A hierarchical navigation point, such as a chapter.
navTarget A target in an auxiliary navigation list; its allowed use depends on the standard edition.
Segment A useful navigable unit that need not be a chapter or heading.
PCM / WAV PCM represents uncompressed audio samples; WAV is a common container for those samples.
Sample position An integer location in audio at a known sample rate, useful for precise internal timing.
AMR-WB+ The extended wideband audio codec required by the cited NLS profile; different from AMR-WB.
3GP The container used for the compressed audio, with NLS-specific constraints.
LUFS A unit for measured perceived loudness; the applicable mastering edition defines the target.
PDTB Protected Digital Talking Book; content protection and authorization for the book/player ecosystem.
MD5 A checksum algorithm used in the specified DTB checksum format. A checksum is not an authorization or authenticity system.
CRC A ZIP integrity check value; the delivery specification constrains how it appears in the headers.
Idempotent job A job whose retry does not create unwanted duplicate effects.
Incremental build Regenerate changed outputs and their dependencies while reusing unaffected accepted artifacts.
S3 AWS object storage; proposed for internal source and artifact storage here.
GPU A processor commonly used to accelerate model inference.
Real-time factor Processing time divided by audio duration under defined hardware and workload conditions.
POP Period of performance: the contractual time window for performing the work.
SOW Statement of work: the document defining requested work and deliverables.
RFI Request for information: a market-research request, distinct from an award or accepted implementation.

For precise format definitions, use J2 or the 2002 standard according to the approved profile. General operational definitions in this table are explanatory additions.

Terms used in the response

An RFI gathers market information; this notice does not award a contract. A capability statement describes what a real supplier can substantiate. A proposed win theme connects a buyer benefit to a method and evidence; it is not proof of superiority. A ROM is an assumption-based rough cost estimate, and its basis of estimate explains the arithmetic. A coverage matrix maps requirements to response and proof. Past performance is attributable completed work, not a proposed design. Government acceptance is the authorized decision, not an internal test pass. Current response context.

Reference

Source library and reading guide

Open the original files and the primary references behind the wiki.

Download this article as MarkdownArticle permalink

Sources were checked for this local edition on October 1, 2026. Attachment citations link directly to the original SAM.gov downloads. Preserved local copies remain available as backups; live links may change. Nearby article citations identify the source of specific claims. Public references do not automatically supersede attachments or resolve procurement ambiguities.

Evidence for the response

The live notice was rechecked October 2, 2026 and is an RFI/Sources Sought. Competitive research uses public provider documentation with explicit limits. No company credentials, prototype results, customer references or approved price are established by this source library. Preserve the four attachments and original analysis as evidence; they are not rewritten to strengthen sales claims.

Original opportunity attachments

Source Open the source What it contributes
Original user analysis Original analysis text Feasibility judgment, architecture proposal, model shortlist and five proof gates. Preserved verbatim; its old sandbox links remain historical text.
Draft SOW Word document · Local backup · Text extract Scope, functions, integration, support, ownership, documentation, phases and milestones. Use section numbers rather than fixed Word page numbers.
J1 1203:2018 construction specification · Local backup 31 pages; version 1.3, January 2019. Exact NLS construction rules.
J2 Z39.86-2005 R2012 · Local backup 95 pages; a different edition from the SOW's stated 2002 target. Its printed page numbers have a 12-page front-matter offset.
J3 QA201801 navigation instructions · Local backup Eight pages; version 2.0, December 2018. How book structure becomes navigation.
Opportunity SAM.gov notice 2027-NLS-0075 Original opportunity and attachment context. Check it for later amendments.

The attachment links above are public SAM.gov download endpoints, without session tokens or API credentials. If a direct download becomes unavailable, return to the opportunity attachment list or use a local backup.

The original analysis and SOW text extract are local supplementary files, not attachments hosted on SAM.gov. The SOW text extract is a convenience copy; layout and table relationships may be flattened, so use the Word original to resolve ambiguities.

PDF links with #page= use one-based file page numbers, not necessarily the printed number on the page. Browser support for opening at a specific PDF page varies, and SAM.gov may download a PDF rather than display it at the requested page.

Standards and production references

Primary reference Why follow it
NLS construction 1203:2025 Compare the newer public construction rules with attached J1.
NLS mastering 1202:2026 Master-audio parameters and heading preparation; clarify application to synthesis.
NLS delivery 1206:2025 Archive layout, ZIP details and Transfer Portal delivery.
NLS specification index Discover the complete applicable specification family. The indexed mastering date differs from the live page.
NLS 1205 protection PDF Catalogued November 2007 reference; direct retrieval failed during this review. Its clauses and governing revision were not independently verified here.
LOC AMR-WB+ format description Additional official format context and link to 1205.
ANSI/NISO 2002 The standard explicitly cited in the SOW and NLS construction profile.
ANSI/NISO 2005 R2012 Online counterpart to the attached later edition.
DAISY PDTB specification Concepts behind protected content and authorization/key material.

Models and reusable software

Reference Scope of evidence
Qwen3 TTS project and 1.7B Base model Documented variants, local loading and exact model-card license.
Chatterbox Official model options, controls and project license.
Kokoro 82M Model card and local usage; example emits 24 kHz audio, making mastering interpretation relevant.
Fish Audio S2 Pro Model capability and licensing pointer; commercial use terms need separate review.
Qwen3 ASR ASR and separate forced-aligner documentation; aligner limit is up to five minutes, not a claim about every ASR mode.
WhisperX Transcription/alignment implementation and documented limitations.
NeMo text processing Normalization and inverse-normalization foundation.
DAISY Pipeline conversion DTBook-to-DAISY3 options; its stated input is 2005-3 DTBook.

These sources document options to evaluate. They do not establish NLS acceptance or benchmark results for this project.

Historical integration evidence and AWS options

Hindenburg's NLS encoder support article and APH Book Wizard's revision history show previous integration with NLS components. They do not confirm current availability, supported operating systems, redistribution terms or cloud deployment.

AWS Batch GPU jobs and Step Functions with Batch document orchestration options. The Library's approved environment and actual workloads determine whether those options fit.

What changed from the original analysis

The wiki reorganizes the analysis into linked topics and adds explanation rather than treating every recommendation as a requirement. Source review added resource-audio timing exceptions, WAV-relative timing semantics, exact attached-edition limits, draft milestones, support/warranty terms, a missing-specifications list, the live mastering edition, and the distinction between S3 staging and the public Transfer Portal.

The original analysis remains available unchanged for comparison. No benchmark, cost quote, accepted architecture, contract interpretation or playback demonstration has been added as a claimed result.

Prototype planning references

The prototype adds an engineering plan, hardware allocations and test targets, all labeled as proposals. Primary sources below document model/runtime capabilities and available hardware, not NLS approval or measured performance. Checked October 1, 2026.

Reference Use in the prototype plan
Qwen CustomVoice Exact proposed preset-voice checkpoint and license label.
MLX-Audio Qwen adapter Community Apple Silicon runtime; conversion needs independent qualification.
Whisper turbo, whisper.cpp, faster-whisper ASR checkpoint and Mac/CUDA runtime paths.
Qwen forced aligner, Qwen 4B instruction model Optional alignment and constrained structural-assistance candidates.
Apple Mac Studio, MacBook Pro M5 Pro/Max Published memory/processor/storage configurations, not performance claims.
AWS G6, AWS G5 Published single-GPU instance specifications; approval and availability still need confirmation.
AWS Pricing Calculator Obtain a region-specific quote; no current price is assumed in the worksheet.