# Prototype implementation plan

**Proposed baseline:** build a working, inspectable system that converts electronic text and supplied commercial audio into sample NLS digital talking books, with an operator able to correct a defect and rebuild the affected output. Start on an Apple Silicon Mac; prove the same workflow on one GPU in an approved AWS development environment. The proposed narration engine is **Qwen3-TTS 1.7B CustomVoice**, with Kokoro-82M as a compact comparison baseline; model selection remains conditional on listening tests and rights review.

**Status: plan only.** The wiki and its hosting are implemented; the production prototype, model benchmarks, purchased hardware, and NLS acceptance are not. Scope, staffing, durations, memory budgets and pass thresholds below are engineering proposals. Contract statements refer to the supplied **draft** SOW, not an awarded or complete contract.

Read [model choices](prototype-models.md), [Mac and AWS hardware](prototype-hardware.md), and [test corpus and acceptance](prototype-tests.md) alongside this plan.

## What the prototype must prove

The prototype should answer four questions: can the chosen voice read complete books faithfully; can the compiler produce a protected book that works on the intended players; can operators resolve exceptions without redoing a whole book; and can the Library operate the delivered system without a required paid speech API or recurring software-access license?

Phase I in the draft calls for an initial working TTS prototype, DTB samples, voice samples, an end-to-end workflow demonstration and a test report. The broader SOW covers both input routes. This proposed prototype deliberately exercises both so commercial-audio risks appear before Phase II. It does not imply that Phase I alone completes all integration, production or support obligations. [Draft SOW §§3.1–3.2, 3.12–3.13; deliverable table §4.1](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/34eacfe92b1a46228b9f90674c4c60e1/download).

## Scope and boundaries

| Area | Included in this prototype | Boundary or follow-on work |
|---|---|---|
| Text input | EPUB, DOCX, RTF and UTF-8 plain text; preserve reading order, headings, source spans and metadata | Confirm whether “Word Documents” includes legacy `.doc`; quarantine unsupported, malformed, encrypted or DRM-protected inputs rather than silently drop content. Scanned-page OCR is a separate extension. |
| Commercial audio | WAV, MP3, OPUS and FLAC fixture coverage; preserve narration, propose navigation from available metadata/text/ASR, synthesize approved announcements | Do not recreate the narrator or invent unavailable page references. Audio-only boundaries need review. |
| Languages and content | English prose first, two approved selectable voices; ordinary headings, lists, notes, numbers and abbreviations | English is a proposed pilot limit, not an SOW exemption. Mixed-language text, equations, poetry, complex tables and illustrations must be flagged and given an agreed reading treatment. Wider language support is a follow-on gate. |
| Output | One agreed audio-plus-navigation profile, WAV masters, required XML/metadata, AMR-WB+, PDTB protection and delivery package | MP3 is optional in the draft. Other multimedia types remain a scoped decision. A WAV/MP3 demo or unprotected package cannot substitute for required protected-output proof. |
| User workflow | Import, inspect structure, select voice, queue job, inspect issues, edit pronunciations/labels, regenerate, compare versions, approve and export | One production team and simple roles; no multi-tenant service, billing or public upload portal. |
| Automation | Three-book batch, durable checkpoints, bounded retries, cancellation, resumable work and dependency-based rebuilds | No promise to sustain the approximately 3,000-book annual background volume until representative benchmarks exist. |
| Deployment | Native Mac model worker, shared API/data contracts, Linux NVIDIA worker in approved AWS test environment | Real Library storage/distribution integration belongs to the agreed interface plan and Phase II acceptance. Prototype adapters may be simulated, visibly labeled. |
| Handover | Source, pinned models and voices where redistribution is permitted, deployment recipe, test evidence and operator/admin instructions | Production HA/DR, large-scale load qualification, full training and support remain later-phase work; backup/restore and access controls still receive prototype tests. |

## Build sequence and stop conditions

<!-- diagram:prototype-gates -->

**Planning allowance: 16 engineering weeks**, assuming roughly four full-time engineering contributors plus fractional project management/specialists, available source rights, an NLS technical contact and timely access to required tooling and players. This is an internal estimate, not a delivery promise. Calendar commitments must follow the actual POP start and Government decisions. The draft sets a **six-calendar-month Phase I deadline**; do not equate that deadline with 24 weeks or assume acceptance/review latency is zero. [Draft schedule](roadmap.md#schedule-in-the-supplied-draft).

| Workstream and proposed window | Concrete work and owner role | Exit evidence / dependency |
|---|---|---|
| Weeks 1–2: baseline and risk probes | Technical lead + audio/ML engineer: requirements matrix, provisional profile, corpus rights, two-voice short sample, protected tiny-book spike, Mac runtime probe | Recorded decisions; encoder/protector/player access requested immediately. A missing codec/protection route blocks a claim of complete DTB success. |
| Weeks 1–4: design package | Technical lead + security/QA contributors: five specified diagrams and eight specified documents | Submit within **30 calendar days of POP start**, regardless of the week labels. Unresolved decisions remain explicit. |
| Weeks 3–6: common core | Backend/DTB engineer: source store, versioned book schema, EPUB/DOCX/RTF/text import, compiler rules and job state machine | Tiny reference book round trip; no unaccounted source text or broken navigation; stable worker contracts. |
| Weeks 5–9: narration and review | ML engineer + product/backend engineer: model adapters, pronunciation overrides, issue queue, transcript/audio comparison, versioned approvals | Full text-origin book, both voices, deliberate pronunciation repair with correct downstream invalidation. |
| Weeks 7–11: commercial path | Backend/DTB + ML engineers: media decoding, transcript alignment, metadata navigation, supplemental speech | Complete commercial-source sample with matching text; audio-only sample with uncertainty surfaced for review. |
| Weeks 10–13: reliability and evaluation | QA/accessibility lead + all engineers: complete corpus, malformed inputs, crash recovery, bounded retries, independent playback | Measured defect and repair rates, all output checks, corpus coverage, accessible operator flow. |
| Weeks 13–16: AWS proof and handover | Technical lead + platform/security contributor: approved AWS sandbox, controlled model transfer, integration rehearsal, clean rebuild and demo | Same reference suite passes; compute/storage/review costs recorded; handover rehearsal and test report. |
| Remaining time before Phase I deadline | Project lead coordinates Government feedback, component delays and corrective work | Preserve contingency explicitly. If required access is late, revise the schedule and report the blocked gate rather than relabel a partial demo as accepted. |

Some work overlaps; format/protection proof, approved output rules and lawful voice/model use constrain later release claims. Model experiments and source-preserving import work can continue while external answers are pending.

## Concrete implementation shape

Use a modular application before distributing every component. **Proposal:** Python API and workers, PostgreSQL for state, an accessible browser operator interface, filesystem artifacts on Mac, and S3 artifacts in AWS. A small job queue and worker leases are sufficient initially; production orchestration can use approved AWS Batch/Step Functions after the basic contracts are stable. See [architecture](architecture.md) and [operations](operations.md).

| Module | Prototype behavior | Durable contract / output |
|---|---|---|
| Importer | Validate file type, sandbox document/media parsers, preserve originals, extract structure and source spans | Source hash, parser version, ordered nodes, warnings, catalog claims with provenance |
| Book service | Version source interpretation, spoken-text decisions, pronunciation rules, voice and output profile | Immutable book version; every accepted passage mapped to a source span or an approved announcement |
| Job service | Enqueue stage jobs, lease work, checkpoint, cancel and resume | Input hash + recipe + stage identity; explicit failed/review/released states; no duplicate delivery |
| Speech worker | Generate one bounded chunk under a book-level voice recipe | PCM audio, duration, model/voice hash, runtime/precision, parameters, peak memory and elapsed time |
| Analysis worker | ASR, alignment, silence/clipping checks and text-difference proposals | Issues tied to exact text/audio versions; no autonomous acceptance |
| Compiler | Assemble accepted masters; derive timing; build XML, encode, protect and package | Profile version, actual delivered files, validator report, checksums and dependency graph |
| Review and delivery | Human resolves exceptions; release requires current evidence; delivery adapter records receipt | Actor, decision, rationale, approved version, package identifier and acknowledgement |

Run local Metal/MLX inference as a native macOS process. Keep database/API containers optional. The Linux CUDA image is a separate build behind the same worker interface, not a claim that a Mac container image has GPU parity. Platform-dependent numerical differences are expected; compare coverage, quality and output validity rather than require identical generated waveforms.

## Operator demonstration

Demonstrate a traceable session: import an EPUB, inspect recovered chapters and metadata, choose an approved voice, generate a chapter, inspect a flagged word, correct its pronunciation, audition before/after, rebuild the affected audio/timing/package, run validators, and navigate the protected result on a target player. Then repeat with a commercial audiobook, preserving its narration and adding only the approved generated material.

The UI needs book/job lists, a navigation-tree editor, text/audio issue review with keyboard controls, voice and lexicon settings, an evidence panel and a release action. A release blocked by stale validation should explain which change invalidated the evidence. Separate operator, reviewer/releaser and administrator permissions even if one person holds several roles during the pilot.

## Deliverables and staffing

Propose four core contributors: a technical lead/platform engineer, a backend/DTB engineer, a speech/ML engineer and a QA/accessibility engineer. Product/UI design, audio editorial expertise, Library integration/security and procurement/license review are fractional support roles; they are real dependencies, not assumed free capacity. Reserve NLS reviewer and player access in the schedule. With one or two engineers, revise scope or time rather than inherit this estimate.

The day-30 design package should map directly to the SOW's five diagrams: **System Content Diagram** (wording in the draft), Component Architecture, Data Flow, DTB Packaging/Output Structure, and Security/Access. Its eight documents cover System Architecture, DTB Integration, Text Processing, Audio Output, Metadata/Packaging, Security/Data Protection, Testing/Validation, and Implementation Schedule. The wiki figures explain the proposal; they do not by themselves constitute all these signed-off deliverables. [SOW §3.13.1 and §4.1](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/34eacfe92b1a46228b9f90674c4c60e1/download).

Map the technical lead to the **Senior Engineer** key-personnel responsibility and allocate a separate **Project Manager**, proposed at 0.25–0.5 FTE during the prototype. The draft identifies both roles in §6.1; actual personnel must satisfy the stated qualifications. The PM owns the decision log, dependency escalation, schedule and status/quality-control reports every 30 calendar days. This management capacity is additional to the four engineering contributors, not hidden in their estimates. [SOW §§3.10–3.11, 4.1, 6.1](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/34eacfe92b1a46228b9f90674c4c60e1/download).

End-of-prototype delivery includes a reproducible code release, dependency/model/voice inventory and license notices, configuration/profile exports, corpus manifest, voice comparison, sample DTBs, test results, benchmark raw data, known limitations, backlog, setup/rollback/restore runbooks, a workflow recording and an operator/admin walkthrough. No source books, private recordings or keys go into the public wiki.

## Decisions before implementation spending

Resolve the output-profile editions, synthetic-audio mastering interpretation, available encoder/protector and OS support, authorized voices, Library account/region/network access, Section H controls and retention, target players and acceptance owner. Obtain throughput and human-review budgets before selecting production capacity. [Decision register](decisions.md).

If an encoder or protection tool is Windows-only, add a controlled Windows worker and a license/deployment review; Mac development does not remove that dependency. If full-book speech quality fails, compare the baseline and challenger or narrow the agreed content class; do not hide failures behind unbounded retries. If a model license cannot meet the ownership/access terms, replace the component before investing in its integration.
