# Proposed local model stack

**Start with Qwen3-TTS-12Hz-1.7B-CustomVoice for narration, and evaluate it against Kokoro-82M.** This is a proposed starting point, not a finding that either model meets NLS requirements. Keep model inference local to the developer Mac or the approved AWS environment; the delivered workflow should not depend on a paid speech API. Training or fine-tuning a new model is outside the initial prototype scope.

## Primary narration choice

The exact upstream artifact is [`Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice`](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice). Its card lists Apache 2.0, preset voices, multilingual synthesis and style instructions. Choose **CustomVoice**, rather than Base, for the first prototype because it provides preset speakers without requiring a new voice-cloning workflow. The official examples use CUDA/BF16. [Official Qwen implementation](https://github.com/QwenLM/Qwen3-TTS).

Start with restrained narration and short sentence-group chunks. Upstream lists Ryan and Aiden as English-native presets, both male. Include other approved English voices in the comparison if the pilot needs a broader selection; multilingual capability alone does not prove equal narration quality for every speaker/language pairing. Treat speaker choice as book-level configuration. Never silently switch engines or voices to repair one sentence.

The [official checkpoint tree](https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice/tree/main) is approximately **4.52 GB** at review time, including associated files. That is download size, not peak inference memory. The “1.7B” name alone cannot determine the total working footprint. Pin the selected revision and record actual memory during the runtime probe.

## Supporting models and alternatives

| Role | Proposed artifact | What it does, and its limit |
|---|---|---|
| Compact TTS comparison | [`hexgrad/Kokoro-82M`](https://huggingface.co/hexgrad/Kokoro-82M) | Small Apache-2.0 model with 24 kHz sample output. Useful for plumbing, announcements and voice comparisons; [English voice catalog](https://huggingface.co/hexgrad/Kokoro-82M/blob/main/VOICES.md) provides more preset choices. Its size does not establish full-book quality. |
| Independent speech check | [`openai/whisper-large-v3-turbo`](https://huggingface.co/openai/whisper-large-v3-turbo) | MIT, 809M-parameter ASR candidate. Transcribes generated and commercial audio to flag discrepancies. Hallucinations and recognition errors mean its transcript is evidence for review, not truth. |
| Alignment experiment | [`Qwen/Qwen3-ForcedAligner-0.6B`](https://huggingface.co/Qwen/Qwen3-ForcedAligner-0.6B) | Apache 2.0; aligns supplied text/audio, with a documented maximum of five minutes per input. Use much shorter overlapping windows and reconcile boundaries. Alignment cannot independently prove the supplied words were spoken. |
| Optional structure assistant | [`Qwen/Qwen3-4B-Instruct-2507`](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507) | Apache 2.0 instruction model for schema-constrained heading/reading-order suggestions on ambiguous input. Use a proposed 4–8K context initially. It must not rewrite narration, invent catalog facts or generate authoritative final XML. |
| Later challenger | [Chatterbox](https://github.com/resemble-ai/chatterbox) | MIT project with several model variants and an embedded watermark. Compare only if the initial two-engine evaluation leaves a meaningful quality gap; validate exact variant, voice rights, runtime and watermark behavior after encoding. |

The prototype needs TTS plus independent audio checking; it does not need every optional model loaded or even integrated on day one. Deterministic parsing remains the first choice for documents with usable structure.

## Mac and AWS runtime paths

| Stage | Apple Silicon development | Linux NVIDIA reference |
|---|---|---|
| Qwen narration | [MLX-Audio Qwen adapter](https://github.com/Blaizzy/mlx-audio/blob/main/mlx_audio/tts/models/qwen3_tts/README.md), beginning with the documented BF16 CustomVoice conversion | Upstream Qwen PyTorch/CUDA runtime; BF16 baseline with pinned dependencies |
| Kokoro comparison | [Upstream Kokoro](https://github.com/hexgrad/kokoro) using CPU or MPS as supported by the pinned version | Same upstream model and code on CUDA/CPU |
| Whisper discrepancy checks | [whisper.cpp](https://github.com/ggml-org/whisper.cpp), Apple Silicon/Metal, with a pinned model conversion | [faster-whisper](https://github.com/SYSTRAN/faster-whisper), with matching model source and recorded decode/precision settings |
| Optional aligner | MLX-Audio conversion only after timestamp checks against known fixtures | Upstream Qwen aligner as reference implementation |
| Optional structure LLM | [MLX-LM](https://github.com/ml-explore/mlx-lm), evaluated 4-bit conversion | Pinned local inference runtime and the same approved source model; benchmark separately |

MLX-Audio is a community implementation. Its support is not a claim that the original CUDA package runs unchanged on Mac. Qualify the conversion on identical passages and on full-book continuity. Metal/MLX uses the Mac's shared memory; a CUDA worker has separate GPU and host memory. [Hardware tiers](prototype-hardware.md).

For the first quality comparison, keep Qwen TTS at BF16. Quantized TTS is a separate, held-out experiment. Record runtime-specific output changes; do not assume a smaller checkpoint preserves fidelity. Use separate inference environments when dependencies conflict.

## Guardrails in the adapter

Each synthesis request carries approved text, language, speaker ID, narration style, pronunciation policy, model and runtime versions, precision and a deterministic job identifier. Save audio duration, sample rate, text hash, elapsed time, memory peak and warnings with the output. Seeded generation improves traceability but does not guarantee identical audio across hardware.

The adapter must reject empty or incomplete output and detect truncation. Kokoro's [pipeline implementation](https://raw.githubusercontent.com/hexgrad/kokoro/main/kokoro/pipeline.py) has bounded phoneme sequences and truncation-related behavior; enforce complete source coverage before calling inference, not merely by reading logs afterward. Use sentence boundaries and source spans to split material, and keep the distinction between synthesis chunks, navigation units and packaged files.

Pronunciation overrides change a separately recorded spoken representation. Do not feed arbitrary SSML to a model unless its adapter explicitly supports it. Check names, numbers, units and negation even when overall ASR word error rate looks good. Use the [test plan](prototype-tests.md) for release evidence.

## Rights, mastering and handover

Apache/MIT labels are useful starting evidence, not blanket clearance of model provenance, voice rights, dependencies, redistribution or the draft's ownership/non-infringement terms. Maintain a reviewed inventory of exact checkpoint revisions, conversions, voice assets, code licenses and any reference-speaker permission. Voice cloning, automatic character voices and style effects are deferred unless explicitly approved for the content.

Pre-fetch approved artifacts into a controlled cache, verify hashes and prove inference with network access disabled. Continued operation must not require a vendor account, hosted model endpoint or license renewal. The final rights arrangement still needs procurement review against [SOW §§2, 3.9, 3.12](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/34eacfe92b1a46228b9f90674c4c60e1/download).

Native synthesis sample rates and token rates are not output-compliance settings. Keep original generated audio, then apply the approved mastering/encoding profile. Upsampling to 44.1 kHz does not settle whether a model satisfies a requirement about initial recording; resolve the [mastering question](decisions.md#specific-mastering-clarification) before making a compliance claim.
