I've been building a personal AI engine that runs entirely on my Mac Studio M4 Max. Text generation was the first step. The natural next question was: can I generate voice locally too? Turns out, yes — and the results are surprisingly good.
The Same Bedtime Page, Three Voices
The newest example is the goodnight page from Susu the Cat, the second of my son’s bedtime books. Same script, three voices — one in the cloud, two on the Mac Studio:
[softly] Goodnight, Susu. <pause> Goodnight Baba. Goodnight Mama. Goodnight Susu, curled up at the end of the bed. <pause> [whispers] Goodnight, Yusuf.
ElevenLabs Eleven v4 — my cloned voice
Cloud — audio tags for delivery — the version that ships in the book — 8.6 s
Kokoro via Voicebox — local
Mac Studio — preset voice “George” — rendered in about 5 seconds — 13.6 s
Qwen Custom Voice via Voicebox — local
Mac Studio — preset voice “Ryan” — rendered in about 20 seconds — 15.0 s
Two things stand out. The obvious one is that only the first sounds like me — the local versions use preset voices, because I haven’t cloned my own voice locally yet. The less obvious one is timing: the local engines read the same page more than 50% slower. That matters for a picture book, where the next page shouldn’t turn before the sentence ends.
What Changed at ElevenLabs: Eleven v4
The bedtime books are narrated with Eleven v4, which is now my default. The big change from the older Multilingual v2 is how you steer it. v2 had dials — stability, style — and you tuned delivery by turning them. v3 and v4 ignore the style dial entirely and take audio tags written into the script instead: [warmly], [softly], [whispers], plus <break time="0.5s" /> for pauses.
One pleasant surprise: on my account, v4 used roughly a tenth of the quota the script length suggested — about 740 credits for an 80-clip, 4,700-character book — and the tags and breaks didn’t appear to count at all. Your plan may differ, so measure before you budget.
Local Voices Inside H_AI: Voicebox
In February I said the next step was building voice into my AI engine. It’s done, through Voicebox, an open-source local voice studio that runs on Apple Silicon through MLX. H_AI talks to it over a local API for both directions: speech-to-text (Whisper) when I record a question on my phone, and text-to-speech when I ask for an answer read aloud.
| Engine | Type | Languages | Best For |
|---|---|---|---|
| Kokoro | 54 preset voices | 8, incl. English | Fast drafts and auditions — a 327 MB model |
| Qwen Custom Voice | 9 preset voices | 10 | Delivery directions written in plain language |
| Qwen, LuxTTS | Voice cloning | 10 / English | A voice from a short reference recording |
| Chatterbox | Voice cloning | 23, incl. Arabic | Arabic — with a reference voice |
| TADA 3B | Voice cloning | 10, incl. Arabic | Multilingual cloning, the heaviest at 8 GB |
The catch for my projects is Arabic. None of the local engines ship an Arabic preset voice; only the cloning engines speak it, so Arabic narration needs a clean 15–30 second recording to clone from — ideally of someone speaking Arabic. In the cloud, I just pick an Arabic voice.
One Script, Two Backends
The newest piece is a batch voice-over tool in H_AI. It takes the same script file I use for ElevenLabs — one entry per page — and renders every page locally through Voicebox, with the same file names. A story player that plays audio/en/p1.mp3 doesn’t care which one made it.
Local engines have never heard of [warmly] and will read it out loud, so the tool strips the tags, splits the script at every <break>, renders each piece, and splices them back together with real silence of the requested length. As the tool’s notes put it: pacing survives the move; emotion does not.
After each batch it checks the results the same way my ElevenLabs pipeline does — flagging clips that are silent, too short, or stretched. Stretched is the clever one: garbled speech still decodes as a perfectly valid file, but it takes far longer per character than its neighbours, so it shows up as an outlier.
Lessons from the Bedtime Stories
Text-to-speech falls apart on bare sound words. My first pass at the animal sounds fed the model strings like Mooooooo with expressiveness turned up, and the output came back warbling. Nonsense gives the model no context. Wrapping each one in a short sentence — “the cow says moo” — and turning expressiveness down fixed it completely.
Write for the ear. Numbers spelled out, no abbreviations, sentences under about 200 characters, one idea per page. Every script goes through a small linter before any audio is generated.
Check the words, not just the file. When I made the local versions of the Susu pages for this article, I ran every clip back through speech-to-text to confirm it said what the script said. The goodnight page came back word-for-word from all three voices — which is why it’s the example above.
The Podcast, Then and Now
This is where the project started. In February I wrote a five-minute podcast script about my Personal AI Engine project and generated it twice — once locally with VibeVoice, once with ElevenLabs using a clone of my voice. In October I ran the exact same script, word for word, through the two new local engines, so all four can be compared directly.
February 2026
Version 1: VibeVoice (Local)
Generated on Mac Studio M4 Max — VibeVoice 7B — completely offline, zero cost
Version 2: ElevenLabs (Cloud)
Generated via ElevenLabs API — cloned voice — cloud-based, paid service
October 2026 — the new local engines
Version 3: Kokoro via Voicebox (Local)
Mac Studio — preset voice “George” — the whole episode rendered in 26 seconds — 5:10
Version 4: Qwen Custom Voice via Voicebox (Local)
Mac Studio — preset voice “Ryan” — rendered in about 9 minutes — 5:34
| Version | Runs | Length | Time to Generate |
|---|---|---|---|
| VibeVoice 7B | Local | 4:31 | A few minutes |
| Kokoro | Local | 5:10 | 26 seconds |
| ElevenLabs (my voice) | Cloud | 5:24 | — |
| Qwen Custom Voice | Local | 5:34 | About 9 minutes |
The new versions were made with the batch voice-over tool described above: the script went in as ten paragraphs, each was rendered separately, and the pieces were joined with a short pause between them. Kokoro is the surprise — a five-minute episode in less time than it takes to listen to the first paragraph. Running every paragraph back through speech-to-text, both new versions matched the script about 98% word for word; nearly all the differences were spelling (“Fast API” for FastAPI), not misreadings.
One thing to keep in mind while listening: the script is from February, so it describes the project as it was then — DeepSeek, Llama and 22 skills. I kept it unchanged on purpose, so every version reads the same words.
What is VibeVoice?
VibeVoice is a frontier long-form conversational TTS model originally developed by Microsoft and now maintained by the open-source community. It generates expressive, multi-speaker audio from plain text — think podcast conversations, narrations, even content with background music that it generates spontaneously based on context. It’s still what I reach for when a script has more than one speaker.
The key innovation is its architecture: continuous speech tokenizers running at 7.5 Hz combined with a next-token diffusion framework. An LLM handles textual context while a diffusion head generates the acoustic details. The result is speech that sounds natural, with proper pacing, intonation, and speaker consistency.
| Model | Max Duration | Speakers | Best For |
|---|---|---|---|
| VibeVoice 0.5B | Real-time | 1 | Low-latency live TTS |
| VibeVoice 1.5B | ~90 min | Up to 4 | Long-form podcasts |
| VibeVoice 7B | ~45 min | Up to 4 | Most realistic output |
Running It Locally
With 64 GB of unified memory on the M4 Max, I can comfortably run the 7B model — the largest and most realistic option. Clone the VibeVoice repo, install dependencies with uv pip install -e ., download the model weights, and run inference from a text file with speaker labels like Speaker 1: Hello, welcome to the show.
Generation takes a few minutes for a full podcast episode. Not real-time, but completely acceptable for producing content. And every bit of it happens on my machine — no data leaves my network, no API calls, no per-minute charges.
Speaker 1: Welcome to The Build Log, where we break down real engineering projects from the ground up. Today, we're looking at something really cool. A developer named Hisham got fed up with cloud AI services, the latency, the costs, the privacy headaches, and decided to build his own personal AI engine that runs entirely on local hardware. Let's dive in.
[ Full script continues for ~19 paragraphs covering hardware, webapp, skills system, security, and roadmap ]
Local vs. Cloud, Eight Months Later
I use ElevenLabs for anything that ships — the bilingual narration in my Islamic Kids Website, the Wireless 101 podcasts, and the bedtime books. Local covers everything before that. Here’s the honest comparison now:
Local (Voicebox + VibeVoice)
- Free and unlimited — ideal for drafts and auditions
- Total privacy — nothing leaves the machine
- Kokoro renders a page in seconds
- Multi-speaker podcasts with VibeVoice
- Built into H_AI for voice in and voice out
- No audio tags — pacing carries over, emotion doesn’t
- Arabic only through voice cloning
- Reads noticeably slower than ElevenLabs
- My own voice isn’t cloned locally yet
ElevenLabs Eleven v4 (Cloud)
- Best quality, and the clone sounds like me
- Audio tags for whispering, warmth, pauses
- 85 languages, with ready Arabic voices
- One recipe worked across two whole books
- Costs credits — though less than expected
- Text is sent to a third party
- The same seed doesn’t reproduce the same audio
- Dependent on an internet connection
What’s Next
In February, the next step was voice input. That’s done — I can record a question on my phone and H_AI transcribes it locally. What’s next now:
Cloning my own voice locally, so local drafts sound like the final cut. Arabic narration locally, through Chatterbox or TADA with a proper Arabic reference recording. And more bedtime books — the series grows with my son, and each new one is another chance to see how far the local voices have come.
The tools are here. The hardware is capable. The open-source community keeps pushing the models forward. Running your own AI voice pipeline is no longer a research project — it’s a weekend project.
This article was written with the assistance of Claude Code.