# Quality checks and operator review

Automation succeeds when defects are detectable and repairs are precise. A generated file is only an intermediate result until content, audio, navigation, format and playback checks agree with the accepted requirements.

The following layered QA design is proposed. The draft requires quality and compatibility testing but does not provide numerical narration-error or review-effort thresholds. Those must be agreed before the pilot is judged. [SOW §§3.5, 3.11, Phase II](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/34eacfe92b1a46228b9f90674c4c60e1/download).

## From a candidate to an accepted release

<!-- diagram:quality -->

## Six layers of correctness

| Layer | Example defect | Proposed evidence |
|---|---|---|
| Source fidelity | A paragraph is omitted or reordered | Source-span coverage and ordering checks against the imported book. |
| Speech fidelity | A number changes or a word is repeated | Independent transcription, protected-token comparison and targeted listening. |
| Acoustic quality | Clipping, an abrupt join, long silence or voice drift | Signal checks plus excerpt and sustained-listening review. |
| Navigation | A label announces one chapter and lands in another | Reference resolution, alignment and interaction tests. |
| Format | Broken XML references, wrong codec, bad checksum or ZIP headers | Schema, semantic, media and package inspectors. |
| Playback | Protection, skip, resume or speed behavior fails on a target player | Recorded test matrix using actual agreed players/apps. |

## ASR evidence is not ground truth

An ASR mismatch can arise from a TTS mistake, a recognizer mistake, or a legitimate difference between written and spoken forms. Forced alignment fits expected text to an audio signal; it cannot independently prove every expected word was spoken.

Combine independent transcription, alignment, acoustic evidence and targeted checks for names, quantities, units, dates and negation. For high-consequence discrepancies, use another recognizer or a human listener. Agreement between models increases evidence but does not guarantee correctness. [WhisperX's documented limitations](https://github.com/m-bain/whisperX) illustrate missing timestamps and other alignment problems.

## A review screen should show enough evidence

For each issue, present the source passage, approved spoken text, recognizer output, short playable excerpt, timeline/waveform, reason for the flag, affected build and repair history. Offer concrete actions: correct pronunciation, edit a proposed boundary, regenerate a passage, or accept a documented exception.

An exception should retain rationale and authorized reviewer identity. Acceptance of one passage should not automatically suppress the same error class throughout a book.

## Defect states and bounded retries

A proposed workflow is **detected → triaged → repair attempted → rechecked → accepted or escalated**. Set an agreed retry budget per issue. Distinguish processing failures, which may be safely retried, from content ambiguities, which often need a decision.

Retest everything invalidated by a repair. Changing a chunk's duration can invalidate later clip times even when the later audio remains unchanged. Tie every check result to the exact artifact hashes and output profile.

## What to measure

| Metric | Definition to agree | Why it helps |
|---|---|---|
| Unresolved content defects | Confirmed omissions, substitutions or repetitions per accepted audio hour | Tracks fidelity more directly than general “naturalness.” |
| Navigation accuracy | Correct labels and destinations divided by evaluated items, by type | Prevents chapter success from masking note/page failures. |
| Review effort | Active human review minutes per accepted audio hour | Measures the practical cost of automation. |
| Repair burden | Attempts and regenerated audio per accepted hour | Exposes unstable models and costly chunk sizes. |
| Throughput | Accepted output hours per elapsed production hour and per compute hour | Separates machine speed from accepted-book delivery. |
| Operational reliability | Failed/resumed jobs, incomplete packages and recovery time | Tests production behavior rather than isolated demos. |

These are proposed definitions, not measured results or contractual targets. Report the test corpus, input class, language, hardware, model version and accepted-output denominator with every result. Avoid hiding rejected books from quality statistics; report rejections separately.

## Playback tests beyond the opening chapter

Exercise all navigation levels, notes/sidebars and skipping behavior, segment stepping, resume, transitions between audio files, and faster playback. Check both protected and unprotected artifacts where relevant. Record the player/app version and profile.

The draft expressly calls for testing with the Library playback environment. The exact device matrix and pass criteria remain [open decisions](decisions.md). Accessibility of the operator interface should also be tested with keyboard and assistive technology as a proposed implementation practice.

## Keep acceptance independent

Use known reference books, deliberate malformed examples, and a held-out evaluation set. An AI coding assistant can help generate tests and analyze regressions, but the acceptance criteria and reference outputs should be reviewed independently from the implementation that produces them.
