A two-host AI podcast generated end to end on my own machine. No cloud voice, no per-minute bill. This page is the running build log.
I've wanted my own podcast for a while, but I didn't want to rent a voice. So I'm building a fully local text-to-speech pipeline for it on the PC I put together for exactly this kind of work. This page is the build log. Each part gets added here as it happens, and I'll keep the earlier parts intact so the mistakes stay visible.
A roughly 30 minute, two-host English podcast about AI, produced entirely on hardware I own. No cloud TTS, no per-minute bill, no terms of service deciding what my show can say. My priorities, in order: it has to sound human, it has to be repeatable, and it has to be easy to tinker with. GPU efficiency is deliberately not on the list.
September 13, 2026. Before installing anything, I picked the candidate models, pinned their versions and weight hashes, and wrote the scoring sheet I'd judge them by. I planned it by running Claude Fable 5.1 and GPT-5.6 Sol on the same brief and feeding their work back into each other until they agreed. The full write-up, including the four reproducibility bugs the cross-review caught before I ran a single command, is on X: I Made Two AIs Argue Before I Installed a Single TTS Model.
RTX 5090 with 32 GB of VRAM, Ryzen 9 7950X, 64 GB DDR5, 2 TB NVMe, Windows 11 Pro. Native Windows, no WSL, which turned out to be a real constraint on which models were even candidates.
Reference voices for the two hosts are synthetic, generated with Kokoro-82M as a neutral third model. I'm intentionally not cloning a real person.
Transcript accuracy with openai-whisper and jiwer, speaker consistency with WavLM-SV, and automated checks for repetition, dropped lines, and abnormal silence. Weights: seconds until it sounds like AI (30%), naturalness (25%), dialogue flow (15%), speaker consistency (15%), artifacts (10%), word error rate (5%), plus hard failure gates. Blind A/B listening is the tiebreaker. Raw and mastered audio are scored separately, and mastering is gain-only loudness normalization so it can't hide anything.
Two models, one brief. Fable 5.1 was the better scout: it found models that fit the use case, surfaced the licensing issue, and pushed for verifying weight hashes before trusting a mirror. GPT-5.6 Sol was the better systems engineer: environment separation, smoke tests, the QC chain, the benchmark structure. Then Fable turned GPT's plan into staged runs with explicit stop points, and GPT became a picky code reviewer once it had something concrete to review. Nothing executes until I've signed off on the previous stage.
Nothing installed yet. Research approved, install plan reviewed through five revisions, moving to execution. Next: machine audit, install, smoke-test both models, a 5 to 7 minute benchmark, blind listening.
October 3, 2026. Both models installed cleanly and passed every automated check. Then I listened, and it sounded like two people reading a script at each other. The full write-up is on X: The Voices Passed. The Conversation Failed.
Before testing whether a control works, measure how much the system moves when you touch nothing. And a quality check only covers what it actually measures. My ears caught the failure none of the metrics did.
Same day update: CosyVoice 3 passed the instruction test that VoxCPM2 failed. Pace, volume, and basic emotion instructions all moved the audio in the right direction, in plain English, well beyond its seed-to-seed noise. Whether "happy" sounds happy or just louder is a listening call I haven't made yet. Next: listen, then put directed voices on the frozen eight-minute script and find out whether they can actually react to each other.