# Speech generation and model evaluation

Use a replaceable text-to-speech (TTS) interface and evaluate several local models on the same material. The names below are candidates from the original analysis. None has been demonstrated here to satisfy NLS acceptance requirements.

## Candidate roles

| Candidate | Documented characteristics | Proposed role |
|---|---|---|
| [Qwen3 TTS](https://github.com/QwenLM/Qwen3-TTS) | Local inference and multiple model/voice modes; the named [1.7B Base weights](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base) list Apache 2.0 | First adapter for narration and approved voice workflows. |
| [Chatterbox](https://github.com/resemble-ai/chatterbox) | Official project lists an English Turbo model and multilingual options under an MIT project license | Independent challenger for consistency and difficult passages. |
| [Kokoro 82M](https://huggingface.co/hexgrad/Kokoro-82M) | Small open-weight TTS model; card lists Apache 2.0 | Compact baseline for narration and announcements. |
| [Fish Audio S2 Pro](https://huggingface.co/fishaudio/s2-pro) | Model-card licensing requires attention, including separate commercial licensing | Conditional evaluation candidate after terms are resolved. |

License labels do not establish rights to every associated voice, dataset, dependency or deployment use. Pin the exact artifacts and record applicable terms in the handover inventory. This wiki's model descriptions are not a legal license clearance.

## A book level narration plan

Before synthesis, select an approved voice, pronunciation lexicon, language handling, and restrained pacing/style settings. Preserve these across chunks. Faithful reading is the default: selectable voices do not imply automatic character acting, added laughs, paraphrasing or sound effects.

Test paragraph or sentence-group chunks. The original analysis suggests experimenting with **20–60 seconds of speech per chunk**; this is a tuning range, not a requirement or measured optimum. Preserve sentence boundaries, and compare voice continuity, pronunciation, joins, and correction cost at different lengths.

## Bounded generation and repair

1. Synthesize a chunk from approved spoken text and a recorded recipe.
2. Transcribe with an independent recognizer and compare meaningful content.
3. Check alignment, clipping, silence, level and joins.
4. Repair a specific discrepancy through a lexicon, text, boundary or generation change.
5. Escalate after an agreed attempt limit rather than retry indefinitely.

Do not silently switch narrators for a difficult sentence. An engine fallback must preserve the approved voice or trigger a deliberate review and potentially a larger regeneration.

## ASR and alignment candidates

[Qwen3 ASR](https://github.com/QwenLM/Qwen3-ASR) provides transcription and a separate forced aligner. [WhisperX](https://github.com/m-bain/whisperX) combines transcription with alignment and documents limitations. Long books need bounded jobs with reconciled overlap, rather than a single whole-book alignment call.

ASR asks what a recognizer hears. Forced alignment asks where expected text fits the audio. These are different checks, and neither establishes perfect narration alone. See [the validation page](validation.md).

## A useful evaluation corpus

Include ordinary prose and long chapters alongside names, dates, quantities, negation, abbreviations, dialogue, notes, lists, tables, mixed-language passages, and repeated terms. Add passages from each supported input class. Keep tuning material separate from a held-out acceptance set.

Measure complete books as well as excerpts. Report unresolved defects per accepted audio hour, review minutes, repair attempts, throughput, and consistency across joins. Do not infer whole-book quality from a 30-second demo or time to first audio.

## Voice governance

Keep approved voice assets, consent or usage rights, version, language scope and allowed transformations with the production recipe. A change to the voice reference can affect every generated passage and therefore should invalidate the appropriate cache and quality evidence.
