# Proposed RFI response draft

**Working draft for review — not submission ready.** This is proposed response language for the current RFI and a starting point for a possible later technical proposal. It has not been sent. Bracketed fields require real bidder information. All implementation statements describe proposed work; no prototype, benchmark, NLS acceptance or corporate qualifications have been demonstrated in this workspace. Verify the [current instructions](opportunity.md) and obtain company approval before submission.

An editable Word copy is available here: [Download the proposed RFI response (.docx)](../drafts/NLS-RFI-Response-Proposed-Draft.docx). This public version contains placeholders and an illustrative pricing scenario only; keep private bidder evidence in the controlled submission copy.

## Response identification

**To:** James Robinson and Brenda Kinlay, Library of Congress Contracts Services  
**Subject:** Response to RFI 2027-NLS-0075 — TTS for DTB Pilot  
**Respondent:** [LEGAL COMPANY NAME]  
**Business address and authorized contact:** [ADDRESS, NAME, TITLE, BUSINESS EMAIL, PHONE]  
**Relevant entity identifiers and business status:** [VERIFIED INFORMATION AS APPLICABLE]  
**Proposed role:** [PRIME / SUBCONTRACTOR / CONFIRMED TEAM AND RESPONSIBILITIES]  
**Response date:** [DATE]

## Executive response

[COMPANY] submits this capability and technical approach for the Library's market research on automated digital talking book production. We propose an inspectable workflow for both electronic text and existing commercial audio, designed around reliable navigation, faithful speech, protected output and continued Library operation.

Our proposed approach would establish the protected-output and target-player path early, then measure full-book fidelity and correction effort before scaling production. The Library would be able to trace source material to spoken content, inspect exceptions, authorize corrections and retain the evidence used to release each book. Model choice would remain replaceable behind stable interfaces.

This response describes a proposed implementation, relevant capabilities to be substantiated below, an illustrative ROM basis for review, and questions that materially affect scope and acceptance. We understand that the current notice is voluntary market research, is not a request for official proposals, provides no reimbursement of response costs and will not itself result in a contract award.

## Bidder capability and relevant experience

[INSERT A FACTUAL COMPANY CAPABILITY PARAGRAPH: what the company currently delivers, its actual role in comparable software/audio/accessibility/AWS work, and evidence supporting each statement. Do not use the proposed architecture as proof of past delivery.]

Provide two or three relevant examples if available; this count is a drafting recommendation, not a stated RFI requirement. For each: [CUSTOMER/PROJECT AND DATES], [COMPANY OR INDIVIDUAL ROLE], [COMPARABLE TECHNICAL/OPERATIONAL CHALLENGE], [VERIFIABLE RESULT], [RELEVANCE TO THIS WORK] and [REFERENCE PERMISSION/CONTACT IN THE CONTROLLED COPY]. Attribute partner and individual experience accurately. If no comparable corporate project exists, disclose that and explain the verified personnel or teaming evidence available.

The proposed key personnel are [PM NAME] and [SENIOR ENGINEER NAME]. [INSERT ACCURATE QUALIFICATIONS AND AVAILABILITY.] We would align these roles with the draft's responsibilities and qualification requirements, including relevant degrees and experience, and provide required resumes with a future proposal. Additional audio QA, accessibility, security and integration support would be supplied through [VERIFIED STAFF OR CONFIRMED PARTNER ARRANGEMENT].

## Proposed technical approach

For electronic text, the proposed system would ingest EPUB, Word, RTF and plain text, preserve original files and reading order, and maintain links from source spans to the approved spoken representation. Deterministic parsing would precede model-assisted suggestions. Uncertain headings, notes, tables or pronunciation would be exposed for review rather than silently rewritten or omitted.

For commercial audiobooks, the workflow would accept WAV, MP3, OPUS and FLAC, preserve supplied narration through approved processing, and use available catalog data, matching text and audio evidence to propose navigation. The system would synthesize only approved announcements and supplemental material. It would not invent unavailable page references or replace the original narrator without an authorized requirement.

Both routes would feed a common versioned book representation. Accepted text/audio and a frozen timeline would drive deterministic XML/metadata generation, approved encoding, PDTB/NLS protection and packaging. The output profile would record governing editions and decisions; outputs would be validated independently and tested on the agreed Library player matrix. Access to approved encoders, protection tools, keys, validators and reference material is an early dependency.

Our proposed first local narration candidate is Qwen3-TTS 1.7B CustomVoice, evaluated against Kokoro-82M with independent Whisper-based discrepancy checks and human listening. This is a candidate selection, not a quality or rights approval. The final engine and voices would be selected from measured fidelity, listener suitability, resource use, supportability and acceptable component rights. We would not assume that upsampling alone resolves synthetic-mastering requirements.

The prototype would use Apple Silicon for authorized development fixtures and a separately qualified Linux/NVIDIA worker in the approved AWS environment. A proposed initial sizing experiment uses 64 GB of Mac unified memory and one AWS L4-class GPU with 24 GB GPU memory. Actual capacity and operating cost would be determined by representative complete-book tests, including failed attempts and human review. These resource figures are unbenchmarked planning inputs.

## Quality and evidence plan

We propose a small complete-book corpus covering both routes, all specified input formats through representative books or fixtures, two selectable approved voices and difficult held-out passages. The pilot would inspect the final encoded/decoded release audio, source coverage, metadata, navigation, mastering and protection. ASR and alignment would flag issues; neither would independently authorize release.

The demonstration would include a tiny protected reference book, a complete text-origin book, a commercial-audio example, a seeded pronunciation/navigation correction, dependent-artifact rebuilds and interrupted-job recovery. Measurements would report accepted audio-hours, defects, correction and reviewer time, compute/memory use and limitations. We propose involving experienced operators and blind or print-disabled readers in evaluating clarity, sustained listening and navigation tasks, with accessible participation and permissions arranged in the project plan.

We would provide the written certifications validating reports required by draft §3.3, with the responsible authorized signatory, exact standards/profile scope and supporting validation evidence identified. We would not certify untested compliance.

All claimed results would identify their corpus, software/profile/model versions, hardware, procedure and acceptance status. We would distinguish internal engineering gates from Government acceptance and retain unresolved defects or blocked tests in the report.

## Delivery management and operational transition

The proposed plan would deliver the five specified architecture views and eight technical documents within 30 calendar days of POP start, followed by Phase I prototype/sample/demo/test artifacts within six calendar months. A 16-week engineering target is an internal planning allowance, subject to staffing and external dependencies; it does not replace the contractual deadline. Phase II integration/operational testing and Phase III deployment would follow the draft's prior-phase acceptance triggers and 90-calendar-day windows. Phase III would include operational ramp-up, production workflow management, performance monitoring and quality control at the agreed scale, as well as deployment and handover. Phase IV would cover ongoing production support, audio/metadata/output validation, security and configuration maintenance, backup/recovery, documentation, reporting and training through the remaining agreed POP. Production volume and human-review responsibility must be defined to resource these obligations.

The PM would own communications, dependencies and 30-day status/QC reporting; the Senior Engineer would own technical integration and evidence quality. Interface agreements, staffing and the complete output profile would be baselined early. The kickoff trigger should be clarified because the draft references both award and POP start.

Handover would include delivered source, build/deployment artifacts, approved model and voice assets where permitted, configuration/profile exports, manuals, training, test records and recovery/rollback procedures. An operator other than the original developer would demonstrate installation and the reference workflow. Support would include the required response within 24 hours and the specified 12-month warranty, with covered defect correction at no additional charge. Detailed support staffing and escalation would be part of the approved service plan.

## Security and rights approach

The design would protect source and generated content through least-privilege access, controlled transfer/storage, separate protection-key management, limited content in logs, bounded egress, backups and verified retention/deletion procedures. Government source materials would be used only for authorized contract performance and disposed of according to the agreed requirements. Section H, approved AWS accounts/regions/services and integration boundaries are required to finalize the controls and cost basis.

We intend to meet the draft's requirement for Library ownership of the resulting TTS software, with no contractor-retained ownership interest and no required separate, recurring or follow-on software license or access purchase as a condition of continued operation. Third-party models, voices, parsers, encoders and protection tools would be inventoried and reviewed for compatible rights and the draft's non-infringement obligation. We request clarification of acceptable treatment of pre-existing/open-source components so that the final rights arrangement is explicit rather than inferred from a model license label.

## Rough order of magnitude

**Drafting note: the following is an illustrative scenario, not an approved estimate from [COMPANY]. Replace or expressly approve it before external use.** At assumed blended selling rates of $150–$200 per hour, 5,200–7,200 hours across Phases I–III, a 15% labor allowance for uncertainty/warranty and $40,000–$110,000 in direct-cost allowances, the example totals **$937,000–$1,766,000**. The [basis of estimate](pricing.md) shows the arithmetic and assumptions.

The phase examples are $503,000–$878,000 for Phase I, $251,500–$490,000 for Phase II and $182,500–$398,000 for Phase III. An illustrative Phase IV unit is $12,000–$32,000 per month for 80–160 labor hours, plus approved incremental infrastructure/direct costs. The illustrative monthly hours are conditional on agreed production volume, review workload, monitoring and coverage; they do not establish capacity for the full ongoing production obligation. The Phase III allowance assumes a limited initial operational ramp and must be recalculated for the agreed production scale. Total support duration is not established. Covered warranty corrections would not be billed again as support.

[REPLACE WITH BIDDER-APPROVED ROM, DATE, CURRENCY, WORK BREAKDOWN, RATE BASIS, ASSUMPTIONS AND EXCLUSIONS.] The estimate should distinguish one-time delivery, operational services and infrastructure, account for Government-furnished resources, and state the effects of content scope, security requirements, tool rights, integration and acceptance. Any RFI figure is nonbinding market-research information, not an official quote.

## Feedback on the draft SOW

We recommend clarifying the following to support comparable estimates and testable acceptance:

1. **Standards baseline:** which editions and precedence govern the draft's 2002 reference, attached 2005 standard, attached construction specification and “current applicable versions”? Please provide the complete applicable specification set and explain “D standards.”
2. **Tooling and playback:** which encoders, protectors, validators, reference DTBs, player devices/apps and key-management interfaces will the Library provide? Please identify supported operating systems, automation interfaces and deployment/redistribution permissions.
3. **Synthetic mastering and scope:** how do initial-recording, heading-audio and mastering rules apply to generated speech? Which languages, multimedia profiles, legacy Word formats and complex content must be accepted?
4. **Security and integration:** please provide Section H, approved AWS boundary/services, content-handling rules, retention/destruction requirements, storage/distribution interface contracts and acceptance responsibilities.
5. **Performance and quality:** what content mix, volume, turnaround, listener criteria and allowable correction effort should inform sizing? We recommend a common corpus, protected-player tasks, source-fidelity checks and accepted-output/review metrics for supplier-neutral comparison.
6. **Rights and delivery terms:** how should Library ownership coexist with approved pre-existing/open-source model and software components? Please clarify total POP, Phase IV scope, warranty/support separation and the kickoff trigger.
7. **Response and future evaluation:** is there a preferred RFI format/length, and what future demonstration evidence would be most useful? We have not assumed unpublished evaluation factors or weights.

## Suggested supporting material

The controlled response package should contain only verified evidence: company capability statement, relevant project examples, personnel summaries/resumes as appropriate, an approved ROM basis, and any actual test artifact with scope/limitations. Planned demonstrations must remain labeled planned. Do not attach the entire internal wiki as a substitute for a focused response.

## Adaptation for a future proposal

If a solicitation is issued, replace this drafting structure with its actual instructions and evaluation order. Prepare an executive value statement supported by evidence, a complete requirement response, technical/quality/security volumes as required, management/staffing and transition plans, verified past performance, and approved pricing/representations. Reconcile every proposed pilot boundary with final scope and identify assumptions or exceptions explicitly. This RFI draft does not itself constitute a compliant future proposal.
