# Prototype corpus and acceptance plan

**These are proposed engineering gates, not Government-approved acceptance thresholds or completed test results.** The prototype should produce evidence for both source routes, full-book speech fidelity, protected playback, repair behavior and operational handover. The draft requires quality checks and functioning Library playback; it does not supply numerical TTS quality or throughput targets. [SOW §§3.5, 3.11, 3.13](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/34eacfe92b1a46228b9f90674c4c60e1/download).

## Corpus with a reason for each item

Use public-domain, created-for-test or explicitly authorized content on developer machines. Record permissions, source hashes, language, duration/word count, expected navigation and known difficult passages. Government-provided materials require the approved handling environment. Proposed corpus: **six complete books, about 30–50 source/output audio-hours in total**, plus a short adversarial suite. The range is a collection target, not a statement about available books or a statistical guarantee.

| Corpus component | Coverage | Human reference work |
|---|---|---|
| Four text-origin books | One each in EPUB, DOCX, RTF and plain text; include prose fiction, nonfiction with notes, numeric/abbreviation-heavy material and weakly structured text | Establish authoritative reading order, metadata, headings and approved spoken forms; distribute difficult features across books. |
| Two commercial-origin books | One with matching text/metadata; one audio-only with limited metadata | Reference chapter boundaries and announcements; explicitly mark unknown page information. |
| Format fixtures | Short valid WAV, MP3, OPUS, FLAC plus all four text formats; corrupt, encrypted and misleading-extension examples | Known success/rejection result for every fixture. Avoid lossy re-encoding merely to enlarge the corpus. |
| Difficult-passage set | At least 100 approved excerpts with names, numbers, dates, abbreviations, notes, repeated words, punctuation, long sentences and language switches | Reviewer-approved pronunciation and reading policy; hold out 20 excerpts from tuning. |
| Tiny protected reference book | Three chapters, nested navigation, announcement, a note and one known page reference | Exact expected metadata/navigation plus target-player checks; use an NLS-approved reference package where available. |

Compare two approved voices and two engines on the same difficult-passage set. Select one recipe before the expensive complete-book run, retain the holdout set, and use a second reviewer to adjudicate disagreements. A six-book corpus demonstrates coverage; it cannot establish a universal error rate for all genres or languages.

## Exit gates

| Gate | Proposed pass condition | Saved evidence |
|---|---|---|
| Source fidelity | Every source span accounted for; every omitted, expanded or normalized spoken passage has an explicit rule or reviewed decision | Source-to-spoken-text map, coverage report, normalization diff; ordinary punctuation transformations recorded by rule |
| Book structure | All required metadata and expected navigation targets are correct in the fixtures; no broken references or unsupported invented page numbers | XML/schema and profile validation, expected-vs-actual navigation comparison |
| Speech quality | Full human listening pass of final encoded/decoded release audio for the six-book corpus, including supplemental announcements; zero known unresolved omissions, hallucinated passages, repetitions or material meaning/pronunciation defects in release candidates | Timestamped defect ledger, repairs and reviewer sign-off; report residual uncertainty and minor defects separately |
| Voice suitability | Blind comparison of clarity, stability, fatigue and joins; NLS reviewers agree voice acceptance threshold before scoring | Per-voice distributions, sample counts, reviewer disagreement and approval; do not claim a generic MOS score proves acceptance |
| Timing and mastering | Measured outputs satisfy the approved profile's clip boundaries, levels, sample rate and encoding rules | Automated measurements plus spot-checked waveform boundaries; [timing explanation](compiler.md#timing-windows-depend-on-clip-type) |
| Encoding and protection | Actual AMR-WB+ plus PDTB/NLS protection using approved tools; output decodes/plays through the authorized path | Tool versions, protected package hashes, validator results, authorized-key playback; never include keys in reports |
| Playback | Every navigation entry in the tiny book and all top-level/representative nested entries in long books work on the agreed NLS hardware/application matrix | Device/application versions, exact actions and observed results; browser audio preview alone is insufficient |
| Repair correctness | Pronunciation, paragraph-length and heading-label changes rebuild every dependent artifact and preserve unaffected accepted chunks | Dependency comparison, new timeline/checksums, old validation invalidated, post-rebuild playback |
| Reliability | Three-book queue survives a worker kill, duplicate request, interrupted upload and restart without lost decisions or duplicate release | Job-event trace, recovered artifacts, deliberate fault results; proposed retry cap: two regeneration attempts after the first |
| Operator accessibility | Complete import/review/repair/release workflow by keyboard and agreed screen-reader combinations; focus, labels, issue announcements and audio controls usable | Manual task test, assistive-technology versions and findings; formal standards baseline agreed with the Library |
| Security and handover | Unauthorized release/content access denied; logs omit source content/secrets; unapproved egress denied; agreed retention/deletion policy verified across originals, intermediates, caches and backups; backup restores a reviewed book; clean installation reproduces the reference workflow | Access, retention/deletion and restore tests, dependency inventory, offline model-loading test and another operator's setup log |

“No known unresolved defect” means the agreed review found and resolved defects; it does not prove the model can never produce an error. Full listening is deliberately expensive for this small pilot so the team can measure what an efficient production review policy would miss. ASR agreement alone must never grant release.

## Benchmark procedure

1. Freeze model, voice, text normalization, chunking, precision, runtime, driver/OS and machine configuration. Record exact model revision hashes; do not compare mutable `latest` labels.
2. Run cold-load and warm inference separately. Use fixed short passages first, then one complete chapter, then whole books. Record output duration and errors along with elapsed time.
3. Measure TTS, ASR/alignment, CPU mastering/encoding/protection, queue wait, transfer and human review separately. Report p50/p95 chunk latency, peak memory, swap, failure and repair rates; include sample counts.
4. Repeat representative runs, including long inputs and interrupted jobs. Count failed generations, discarded audio and repair attempts in cost, rather than measuring only accepted best runs.
5. Repeat the same corpus on Mac and the chosen AWS instance. A faster machine is useful only if accepted output and repair effort remain acceptable. Qualify quantized and converted model artifacts separately.

**RTF (real-time factor) = processing seconds ÷ resulting audio seconds** for a stated stage. RTF 0.5 means that stage takes half as long as its audio; RTF 2 takes twice as long. Do not use time-to-first-audio or streaming latency as whole-book throughput. The [hardware worksheet](prototype-hardware.md#estimate-compute-and-review-effort) is a planning tool until measured inputs replace its examples.

Report accepted audio-hours per paid instance-hour, cost per accepted audio-hour, and reviewer minutes per accepted audio-hour. Suggested feasibility target for discussion: combined automated processing at or below the audio duration on the selected cloud worker, with no sustained memory pressure. This is a proposed test target, not a measured result or an NLS requirement; failure should trigger a sizing/quality tradeoff, not suppression of data.

## Traceability and report format

Each requirement row should contain an ID, exact SOW/specification edition and section, implementation version, test fixture, expected result, observed result, evidence URI, defect ID and acceptance authority. Status values should distinguish **not tested**, **passed**, **failed**, **blocked** and **not applicable with rationale**. A missing player or protector is blocked, not passed.

The final report includes corpus coverage, raw measurements, voice comparisons, issue severity definitions, resolved and open defects, security/environment exceptions, playback evidence and the next-phase recommendation. Record both successes and excluded material. Acceptance should be attributed to the authorized reviewer; the engineering team's internal gates do not replace Government acceptance.

## Demonstration agenda

Allow about 60–90 minutes as a planning assumption: inspect the corpus/profile and rights; create a short text-origin book; audition two voices; review and repair a flagged passage; inspect regenerated navigation and hashes; show the commercial-audio path; navigate protected outputs on target hardware; inject/recover one worker failure; and review benchmark and handover evidence. Pre-generated complete books complement the live short example so the demonstration does not mistake a polished excerpt for long-form validation.
