# Mac hardware and contract-aligned cloud options

**Proposed starting configuration:** an actively cooled Apple Silicon Mac with **64 GB unified memory and a 1 TB SSD**, plus a short-lived **AWS g6.2xlarge** test worker with **one NVIDIA L4, 24 GB GPU memory, 8 vCPUs and 32 GiB host RAM**. Run model stages sequentially at first. This is a practical provisioning proposal for the [model stack](prototype-models.md), not a measured minimum, throughput guarantee or purchase order. [AWS G6 specifications](https://aws.amazon.com/ec2/instance-types/g6/).

## What the draft establishes about cloud providers

The supplied draft says to integrate with NLS AWS systems (§3.7), and describes existing S3-based storage (§1). It does **not** provide a cloud-provider allowlist or explicitly require every workload to execute exclusively in AWS. Because Section H is absent, the permitted account, region, services, data residency, network boundary and security authorization cannot be determined from these attachments alone. [Draft SOW §§1, 3.6–3.7, 5.1](https://sam.gov/api/prod/opps/v3/opportunities/resources/files/34eacfe92b1a46228b9f90674c4c60e1/download).

| Environment / provider | Evidence of permission | Proposed use |
|---|---|---|
| Library-approved AWS environment | AWS integration is explicit; particular compute service/account/region approval is not established | Primary candidate for cloud inference, storage and integration testing. Confirm service availability and GPU quota before booking work. |
| AWS GovCloud | Not named in the supplied draft | A conditional AWS option only if the Library specifies that boundary; do not assume it is required, sufficient or interchangeable with commercial AWS. |
| Azure, Google Cloud, Oracle Cloud | No authorization established by the supplied attachments | Do not include as approved runtime choices. Request a written determination before any contract-content processing or deployment there. |
| RunPod, Lambda, CoreWeave or other GPU hosts | No authorization established by the supplied attachments | Potential technical alternatives only after provider and data-handling approval; excluded from the baseline. |
| Developer Mac | Proposed engineering environment, not an expressly approved place for Government content | Public-domain, synthetic or explicitly authorized fixtures until content-handling permission is confirmed. |
| This Cloudflare Pages wiki | Public documentation hosting already requested by the user | Documentation only. Its deployment establishes no permission to host source books, protection keys, private recordings or inference there. |

Cloud certifications alone would not establish contractual permission. Keep software-access rights and consumption-based infrastructure costs as separate questions; the draft's non-subscription language should be clarified for the final procurement arrangement.

## Mac development tiers

| Tier | Proposed hardware | Intended use and tradeoff |
|---|---|---|
| Reuse existing hardware | Apple Silicon, 16 GB unified memory; adequate free disk | API/UI/compiler work, Kokoro and short serialized smoke tests. Useful starting machine; not the buying target for sustained Qwen plus QA development. |
| Practical sequential-worker floor | Apple Silicon Pro/Max class, 32–36 GB unified memory, 1 TB SSD | One model loaded at a time, bounded chunks/context, few competing apps. Verify actual peak memory and swap before accepting it for whole-book runs. |
| Recommended desktop | **Mac Studio, M4 Max, 16-core CPU / 40-core GPU, 64 GB unified memory, 1 TB SSD** | Comfortable development headroom for speech experiments and application tooling. Apple lists this memory/storage combination; the recommendation is an engineering judgment, not an Apple TTS benchmark. |
| Portable alternative | **MacBook Pro, M5 Pro with 20-core GPU, 64 GB unified memory, 1 TB SSD**; M5 Max optional | Similar memory budget in a portable, actively cooled form. More GPU cores may help, but throughput must be measured on the selected configuration. |
| Larger workstation | 96–128 GB unified memory | Defer unless simultaneous model workers, a larger LLM or training becomes an evidenced need. Extra capacity is not required by this proposed scope. |

Apple's published configurations: [Mac Studio specifications](https://support.apple.com/en-us/122211) and [M5 Pro/Max MacBook Pro specifications](https://support.apple.com/en-us/126318). These establish supported configurations, not current inventory, delivery dates or prices. Confirm availability before purchase. An existing 48 GB machine is also a reasonable evaluation workstation; do not replace it before the memory/throughput probe.

The Mac's memory is shared by GPU, OS and applications; it is not equivalent to the same number of gigabytes of dedicated NVIDIA VRAM. Reserve at least **200 GB initially free** for the pilot, plus a backup destination. Use native Metal/MLX workers; validate Linux/CUDA separately. Headphones for careful listening, representative NLS playback hardware and an accessible test browser/screen reader are as important to acceptance as the GPU.

## Working-memory budget to validate

| Process | Proposed starting allocation | Scheduling policy |
|---|---|---|
| Qwen 1.7B BF16 TTS | 8–12 GB | Batch 1, bounded sentence groups; measure cache/codec/framework peaks |
| Kokoro baseline | 2–4 GB | Load instead of Qwen for a comparison run |
| Whisper turbo ASR | 3–6 GB | Inspect accepted/generated chunks after synthesis; no required co-residency |
| Optional forced aligner | 2–4 GB | Short windows; reconcile boundaries |
| Optional 4B LLM at 4-bit, 4–8K context | 4–8 GB | Separate structural suggestion job; no whole-book context assumption |
| OS, database, API, browser and build tools | Reserve roughly 10–16 GB on Mac | Limit container memory and competing apps; watch memory pressure |

**All allocations above are estimates to test, not vendor minimum specifications.** They depend on runtime, precision, input length, batches and conversion. Do not add every row as if all models must reside simultaneously. Conversely, do not use raw checkpoint size as total RAM. Reduce batch/context first, then serialize or unload workers before buying larger hardware. [Model evidence and runtime choices](prototype-models.md).

## AWS instances and surrounding resources

| Candidate | Published hardware | Prototype role |
|---|---|---|
| **g6.2xlarge — first choice to test** | 1 × L4, 24 GB GPU memory; 8 vCPUs, 32 GiB RAM | One sequential inference worker; separate CPU compilation if host memory becomes tight |
| g6.4xlarge | Same single L4 / 24 GB GPU memory; 16 vCPUs, 64 GiB RAM | More host-side parsing/media/cache headroom; it does not double VRAM |
| g5.2xlarge — availability fallback | 1 × A10G, 24 GB GPU memory; 8 vCPUs, 32 GiB RAM | Benchmark against G6 if approved region/quota or quoted cost makes it preferable |
| g5.4xlarge | Same single A10G / 24 GB GPU memory; 16 vCPUs, 64 GiB RAM | Host-memory alternative; validate the exact backend independently |

Hardware values come from [AWS G6](https://aws.amazon.com/ec2/instance-types/g6/) and [AWS G5](https://aws.amazon.com/ec2/instance-types/g5/) product tables; GB and GiB retain AWS's published units. Family support does not establish capacity, regional availability or Library approval. Avoid fractional-GPU instances for the first reproducibility baseline. No multi-GPU/H100 reservation is justified by current evidence.

**Proposed supporting resources:** one 4-vCPU/16-GiB CPU application/compile node (or equivalent approved service), PostgreSQL, a private artifact bucket, a private model cache, encrypted 250 GB scratch/EBS capacity for the GPU node, bounded log retention, and backup storage. The initial CPU/database services may share a development node; isolate them for later production requirements. Use durable storage for checkpoints; instance-local scratch is replaceable. Start with on-demand test windows and shut workers down when idle; test interruption recovery before considering Spot.

## Same book contract, different execution environments

<!-- diagram:prototype-environments -->

Package identical schema versions, text fixtures, voice recipes and output profiles for both platforms. Compare semantic fidelity, timing validity, memory and accepted throughput. Do not require byte-identical stochastic narration across Metal and CUDA. Test the approved encoder/protector on its actual supported OS; if it is Windows-only, budget a separate approved worker rather than assume Linux or macOS support.

## Estimate compute and review effort

<!-- diagram:prototype-sizing -->

The worksheet uses explicit assumptions; its defaults are illustrative, not benchmarks. It models one GPU per instance and one selected GPU stage at a time. Generated hours include full text narration and supplemental commercial-audio speech. The generation multiplier covers additional synthesis and checking of replacement audio. Commercial original audio is checked once; repeated full-source analysis, optional LLM work, CPU encoding, queue delays and storage/transfer need additional budget. Effective utilization converts busy processing hours to allocated instance-hours; aggregate wall time divides by worker count, while total paid capacity remains unchanged in this simplified model.

Human review is separate and may dominate elapsed delivery. Pricing is deliberately entered by the reader because region, operating system, purchasing arrangement and date change it. Use a current [AWS Pricing Calculator estimate](https://calculator.aws/) for the approved region, and add CPU/DB, EBS/S3, backups, logs, network egress, player access, engineering and human-review costs. Set an account budget and worker time limit before any cloud experiment; no cloud instances are provisioned by this plan.

## Storage sizing and buying decision

As an arithmetic example, **44.1 kHz, 24-bit mono PCM uses about 0.476 GB per hour**: sample rate × bytes/sample × channels × seconds. Fifty hours with four retained PCM versions is about 95 GB before originals, models, intermediates and backups. This example is not an approved mastering profile. Stereo doubles that quantity; compression and retention change actual needs.

Buy or reserve only after a 30-minute representative narration/analysis run, a long-chapter run and the protected-output spike establish fit. Capture peak memory, sustained swap, per-stage RTF, reviewer effort and tool compatibility. Scale worker count only after a three-book batch and fault recovery pass. Approve production capacity against accepted audio-hours and required turnaround, not model parameter count alone.
