Generating Voice from Text Using Local Models

Running open-source text-to-speech locally on a Mac Studio M4 Max — now built into my AI engine — and comparing it with ElevenLabs Eleven v4 on real bedtime stories

I've been building a personal AI engine that runs entirely on my Mac Studio M4 Max. Text generation was the first step. The natural next question was: can I generate voice locally too? Turns out, yes — and the results are surprisingly good.

Update, October 2026. When I first wrote this in February, local voice meant one model, VibeVoice, and one experiment. Since then voice has become part of everyday projects: bedtime stories for my son narrated in my own voice, a set of local voice engines wired into H_AI through Voicebox, and a batch tool that renders a whole story locally or through ElevenLabs from the same script. ElevenLabs moved too — my default is now Eleven v4. New examples and lessons below; the original February comparison is still here.

The Same Bedtime Page, Three Voices

The newest example is the goodnight page from Susu the Cat, the second of my son’s bedtime books. Same script, three voices — one in the cloud, two on the Mac Studio:

[softly] Goodnight, Susu. <pause> Goodnight Baba. Goodnight Mama. Goodnight Susu, curled up at the end of the bed. <pause> [whispers] Goodnight, Yusuf.

ElevenLabs Eleven v4 — my cloned voice

Cloud — audio tags for delivery — the version that ships in the book — 8.6 s

Kokoro via Voicebox — local

Mac Studio — preset voice “George” — rendered in about 5 seconds — 13.6 s

Qwen Custom Voice via Voicebox — local

Mac Studio — preset voice “Ryan” — rendered in about 20 seconds — 15.0 s

Two things stand out. The obvious one is that only the first sounds like me — the local versions use preset voices, because I haven’t cloned my own voice locally yet. The less obvious one is timing: the local engines read the same page more than 50% slower. That matters for a picture book, where the next page shouldn’t turn before the sentence ends.

What Changed at ElevenLabs: Eleven v4

The bedtime books are narrated with Eleven v4, which is now my default. The big change from the older Multilingual v2 is how you steer it. v2 had dials — stability, style — and you tuned delivery by turning them. v3 and v4 ignore the style dial entirely and take audio tags written into the script instead: [warmly], [softly], [whispers], plus <break time="0.5s" /> for pauses.

🏷️
Tags, not dials Expression lives in the script. Settings that worked on v2 silently do nothing on v4
✅
One recipe, 106 clips Stability 0.5, similarity 0.8, with tags — two bilingual books, first try, nothing re-cut
🌍
85 languages Arabic narration with a choice of voices, which still matters a lot for my projects
🎲
Not reproducible The same seed and text gave different audio twice. Keep the file you like — you can’t regenerate it

One pleasant surprise: on my account, v4 used roughly a tenth of the quota the script length suggested — about 740 credits for an 80-clip, 4,700-character book — and the tags and breaks didn’t appear to count at all. Your plan may differ, so measure before you budget.

Local Voices Inside H_AI: Voicebox

In February I said the next step was building voice into my AI engine. It’s done, through Voicebox, an open-source local voice studio that runs on Apple Silicon through MLX. H_AI talks to it over a local API for both directions: speech-to-text (Whisper) when I record a question on my phone, and text-to-speech when I ask for an answer read aloud.

Engine Type Languages Best For
Kokoro 54 preset voices 8, incl. English Fast drafts and auditions — a 327 MB model
Qwen Custom Voice 9 preset voices 10 Delivery directions written in plain language
Qwen, LuxTTS Voice cloning 10 / English A voice from a short reference recording
Chatterbox Voice cloning 23, incl. Arabic Arabic — with a reference voice
TADA 3B Voice cloning 10, incl. Arabic Multilingual cloning, the heaviest at 8 GB

The catch for my projects is Arabic. None of the local engines ship an Arabic preset voice; only the cloning engines speak it, so Arabic narration needs a clean 15–30 second recording to clone from — ideally of someone speaking Arabic. In the cloud, I just pick an Arabic voice.

One Script, Two Backends

The newest piece is a batch voice-over tool in H_AI. It takes the same script file I use for ElevenLabs — one entry per page — and renders every page locally through Voicebox, with the same file names. A story player that plays audio/en/p1.mp3 doesn’t care which one made it.

Local engines have never heard of [warmly] and will read it out loud, so the tool strips the tags, splits the script at every <break>, renders each piece, and splices them back together with real silence of the requested length. As the tool’s notes put it: pacing survives the move; emotion does not.

After each batch it checks the results the same way my ElevenLabs pipeline does — flagging clips that are silent, too short, or stretched. Stretched is the clever one: garbled speech still decodes as a perfectly valid file, but it takes far longer per character than its neighbours, so it shows up as an outlier.

Draft and iterate locally for free, then spend cloud credits on the final cut. Same script, one command to switch.

Lessons from the Bedtime Stories

Text-to-speech falls apart on bare sound words. My first pass at the animal sounds fed the model strings like Mooooooo with expressiveness turned up, and the output came back warbling. Nonsense gives the model no context. Wrapping each one in a short sentence — “the cow says moo” — and turning expressiveness down fixed it completely.

Write for the ear. Numbers spelled out, no abbreviations, sentences under about 200 characters, one idea per page. Every script goes through a small linter before any audio is generated.

Check the words, not just the file. When I made the local versions of the Susu pages for this article, I ran every clip back through speech-to-text to confirm it said what the script said. The goodnight page came back word-for-word from all three voices — which is why it’s the example above.

The Podcast, Then and Now

This is where the project started. In February I wrote a five-minute podcast script about my Personal AI Engine project and generated it twice — once locally with VibeVoice, once with ElevenLabs using a clone of my voice. In October I ran the exact same script, word for word, through the two new local engines, so all four can be compared directly.

February 2026

Version 1: VibeVoice (Local)

Generated on Mac Studio M4 Max — VibeVoice 7B — completely offline, zero cost

Version 2: ElevenLabs (Cloud)

Generated via ElevenLabs API — cloned voice — cloud-based, paid service

October 2026 — the new local engines

Version 3: Kokoro via Voicebox (Local)

Mac Studio — preset voice “George” — the whole episode rendered in 26 seconds — 5:10

Version 4: Qwen Custom Voice via Voicebox (Local)

Mac Studio — preset voice “Ryan” — rendered in about 9 minutes — 5:34

Version Runs Length Time to Generate
VibeVoice 7B Local 4:31 A few minutes
Kokoro Local 5:10 26 seconds
ElevenLabs (my voice) Cloud 5:24 —
Qwen Custom Voice Local 5:34 About 9 minutes

The new versions were made with the batch voice-over tool described above: the script went in as ten paragraphs, each was rendered separately, and the pieces were joined with a short pause between them. Kokoro is the surprise — a five-minute episode in less time than it takes to listen to the first paragraph. Running every paragraph back through speech-to-text, both new versions matched the script about 98% word for word; nearly all the differences were spelling (“Fast API” for FastAPI), not misreadings.

One thing to keep in mind while listening: the script is from February, so it describes the project as it was then — DeepSeek, Llama and 22 skills. I kept it unchanged on purpose, so every version reads the same words.

What is VibeVoice?

VibeVoice is a frontier long-form conversational TTS model originally developed by Microsoft and now maintained by the open-source community. It generates expressive, multi-speaker audio from plain text — think podcast conversations, narrations, even content with background music that it generates spontaneously based on context. It’s still what I reach for when a script has more than one speaker.

The key innovation is its architecture: continuous speech tokenizers running at 7.5 Hz combined with a next-token diffusion framework. An LLM handles textual context while a diffusion head generates the acoustic details. The result is speech that sounds natural, with proper pacing, intonation, and speaker consistency.

Model Max Duration Speakers Best For
VibeVoice 0.5B Real-time 1 Low-latency live TTS
VibeVoice 1.5B ~90 min Up to 4 Long-form podcasts
VibeVoice 7B ~45 min Up to 4 Most realistic output

Running It Locally

With 64 GB of unified memory on the M4 Max, I can comfortably run the 7B model — the largest and most realistic option. Clone the VibeVoice repo, install dependencies with uv pip install -e ., download the model weights, and run inference from a text file with speaker labels like Speaker 1: Hello, welcome to the show.

Generation takes a few minutes for a full podcast episode. Not real-time, but completely acceptable for producing content. And every bit of it happens on my machine — no data leaves my network, no API calls, no per-minute charges.

Speaker 1: Welcome to The Build Log, where we break down real engineering projects from the ground up. Today, we're looking at something really cool. A developer named Hisham got fed up with cloud AI services, the latency, the costs, the privacy headaches, and decided to build his own personal AI engine that runs entirely on local hardware. Let's dive in.

[ Full script continues for ~19 paragraphs covering hardware, webapp, skills system, security, and roadmap ]

Local vs. Cloud, Eight Months Later

I use ElevenLabs for anything that ships — the bilingual narration in my Islamic Kids Website, the Wireless 101 podcasts, and the bedtime books. Local covers everything before that. Here’s the honest comparison now:

Local (Voicebox + VibeVoice)

  • Free and unlimited — ideal for drafts and auditions
  • Total privacy — nothing leaves the machine
  • Kokoro renders a page in seconds
  • Multi-speaker podcasts with VibeVoice
  • Built into H_AI for voice in and voice out
  • No audio tags — pacing carries over, emotion doesn’t
  • Arabic only through voice cloning
  • Reads noticeably slower than ElevenLabs
  • My own voice isn’t cloned locally yet

ElevenLabs Eleven v4 (Cloud)

  • Best quality, and the clone sounds like me
  • Audio tags for whispering, warmth, pauses
  • 85 languages, with ready Arabic voices
  • One recipe worked across two whole books
  • Costs credits — though less than expected
  • Text is sent to a third party
  • The same seed doesn’t reproduce the same audio
  • Dependent on an internet connection
The gap between local and cloud TTS keeps closing. Local is now good enough to draft a whole book for free; the cloud still earns its place on the final cut — especially when the voice has to be mine.

What’s Next

In February, the next step was voice input. That’s done — I can record a question on my phone and H_AI transcribes it locally. What’s next now:

Cloning my own voice locally, so local drafts sound like the final cut. Arabic narration locally, through Chatterbox or TADA with a proper Arabic reference recording. And more bedtime books — the series grows with my son, and each new one is another chance to see how far the local voices have come.

The tools are here. The hardware is capable. The open-source community keeps pushing the models forward. Running your own AI voice pipeline is no longer a research project — it’s a weekend project.

This article was written with the assistance of Claude Code.