Skip to content
CSuite
Use caseAudioCreatorsAugust 2, 20269 min read

AI voiceovers without a studio: podcasts, videos, and audiobooks

We made this post's podcast intro from a paragraph of text: nine cents, seven seconds. Listen first, then the one rule you can't break.

Listen before you read
Nobody recorded this podcast intro. It was typed.
Generated with MiniMax Speech 2.8 through Runware’s API on August 2, 2026: first result kept, no rerolls, no editing. 414 characters of text in, 25 seconds of audio out, 6.8 seconds of waiting, $0.04 billed. Every clip in this post is a real generation, and every receipt is shown.

If you skipped the player above, go back and press play. Those 25 seconds were never spoken aloud. There was no microphone, no booth, no editing pass to clean up a fluffed line. The pauses, the emphasis, even the small breath you can hear between sentences: all of it was invented by a model from a paragraph of text, in less than seven seconds, for four cents.

For most of computing history, text-to-speech announced itself within one word. It was the voice of GPS wrong turns and hold-music menus, useful and unmistakably fake. Somewhere in the last two years it stopped announcing itself, and most people outside the industry haven’t noticed yet. This post walks the whole distance: what changed, what a real voiceover session costs now (we kept every receipt), how you direct a read you can’t perform yourself, what that unlocks for podcasts, videos, and audiobooks, and the one rule about other people’s voices that is not optional. It’s part of the same everyday-abilities story as our plain-English tour of what AI can actually do, zoomed all the way in on the speaking part.

The robotic voice you remember is gone

The fastest way to hear a decade of progress is to hold the two eras next to each other. Below, the same 152 characters read twice. The first voice is Fred, the classic Mac text-to-speech voice, which still ships with every Mac today and sounds essentially the way desktop computers have sounded since the nineties. The second is MiniMax Speech 2.8, a current generation model, given the identical text and nothing else.

Then and now: the same two sentences
1990s technology: Fred
macOS's built-in voice, generated free with the `say` command.
2026 technology: MiniMax Speech 2.8
Same text, voice preset English_Trustworth_Man, $0.015 billed.
Both clips generated on August 2, 2026, first take each, no edits. The text: a two-sentence evening weather sign-off.

The difference isn’t polish. It’s a different technology. Fred’s generation of systems assembled speech from rules and recorded fragments, which is why every sentence has the same shape. Modern voice models are generative: trained on enormous amounts of real speech, they predict audio the way a chatbot predicts words. Nobody programmed the pause before a clause or the falling pitch at a sign-off. The models absorbed those habits from people, so their output inherits the texture of people, breath and hesitation included.

This post’s whole recording session cost nine cents

Claims about cheap AI audio usually stay vague, so here is our actual session, itemized. Every clip embedded in this post came from one sitting on the morning of August 2, 2026, through Runware’s MiniMax Speech 2.8 endpoint, which bills $0.10 per 1,000 characters of input at the HD tier. We kept the first result of every request. No rerolls, no cherry-picking, no audio editing of any kind.

The session, itemized · MiniMax Speech 2.8 via Runware · August 2, 2026
ClipVoice presetTextAudioWaitBilled
Podcast introEnglish_expressive_narrator414 chars25.0s6.8s$0.041
Weather sign-offEnglish_Trustworth_Man152 chars7.9s4.5s$0.015
Guest intro, plainEnglish_CaptivatingStoryteller76 chars5.2s3.7s$0.008
Guest intro, directedEnglish_CaptivatingStoryteller77 chars7.0s5.4s$0.008
The laughing ad readEnglish_PlayfulGirl164 chars8.9s4.7s$0.016
Five clips, 54 seconds of finished audio, $0.088 total. Billed amounts are the API’s own includeCost figures, not estimates.

The arithmetic generalizes nicely: at a normal narration pace, 1,000 characters is about a minute of finished audio, so a dime buys a minute and six dollars buys an hour. A ten-minute weekly podcast intro habit costs about a dollar a year. These are list prices from a metered API, the pay-per-use route; subscription apps bundle the same capability differently, and we’ll come back to what they charge in the closing section.

The whole workflow: a paragraph goes in, a voice comes out. Illustration generated with Seedream 4.5 via Runware.

Speed deserves its own sentence. The 25-second intro took 6.8 seconds to generate. Every clip came back faster than it plays. Against the traditional path (write the script, book the voice actor or clear your throat, record, fix the flubs, master the levels), the change isn’t a discount. It’s a different category of activity, the way typing a document is a different activity from typesetting it.

Typing is the new directing

A voice you didn’t perform still needs direction, and the surprise is where the directing happens: in the text itself. You cast by choosing a preset (the model we used ships 332 voices; the one in our hero clip is literally named English_expressive_narrator). You pace with punctuation. You fix pronunciation with spelling. The script is the control panel.

We ran the classic stress test: an Irish name most humans misread. The plain sentence below uses the standard spelling, Siobhan. In the directed version we respelled it the way it sounds (Shiv-awn) and added an ellipsis where a host would let the name land. Same sentence, same voice, one authorial pass.

The same sentence, plain and directed
Take one: as written
“Our guest tonight is Siobhan Nguyen, whose debut novel comes out on Tuesday.”
Take two: respelled and paced
“Our guest tonight is Shiv-awn Win... whose debut novel comes out on Tuesday.”
One honest surprise: we expected take one to stumble on the name, and a speech-to-text round-trip suggests the model handled Siobhan correctly on its own. The techniques still matter: the ellipsis alone stretched an identical sentence from 5.2 to 7.0 seconds, a real pause you can hear, and respelling remains the dependable fix for names the model does miss.

There are also proper dials when text isn’t enough: a speed parameter (0.5x to 2x on the model we used), language hints for mixed-language scripts, and emotion controls on some models. But in practice, most direction is writing. Read your script aloud once, notice where you naturally pause or lean, and put that on the page as punctuation and word choice. Voice work turns out to be an editing skill, which is good news for anyone who already writes.

The podcast, the video, the audiobook: real jobs now

Three jobs used to have a studio, a budget, or a voice actor standing between the words and the audience. All three are now a text box.

The classic home setup: mic, headphones, editing software, and many takes. The new alternative skips all four, for better and worse. Photo by Soundtrap on Unsplash.

Podcasts and video narration. Intros, outros, ad reads, and full narration for explainer videos are the natural first projects: short, scripted, and re-recordable by retyping. The workflow difference compounds over time. When episode 40 needs the sponsor name changed in the intro, you edit a sentence and regenerate for a nickel instead of matching mic distance and room tone to a six-month-old session. Small businesses hit this from the other side: the promo video that was never worth a voiceover budget suddenly is, a pattern our one-person marketing month priced out in a different modality.

Audiobook listening keeps growing because it fits the commute. The catalog is now growing faster than human narrators can read. Photo by Andriyko Podilnyk on Unsplash.

Audiobooks. This is where the shift stopped being hypothetical. US audiobook revenue reached $2.43 billion in 2025, up 9%, and publishers reported over 750,000 active titles, a 43% jump in a single year. A quarter-million new audiobooks did not come from booking a quarter-million narrators. Audible’s Virtual Voice program lets self-published authors add synthetic narration, labels every result “Narrator: Virtual Voice,” and by mid-2025 had passed 50,000 titles and expanded to publishers with over 100 voices in four languages. Run our dime-a-minute math on a novel: an 80,000-word book is roughly 480,000 characters, call it $50 of synthesis at list price, against the thousands a professional narration run costs. That gap is why the catalog exploded. Whether a synthetic voice can hold a listener for nine hours is a genuinely different question, which is why our audiobook-narration guide tests voices on stamina rather than demos.

Where the seams still show

An honest tour includes the stress test that fails. We asked for the hardest small thing we know: a host cracking up mid-sentence and recovering, the kind of moment that makes radio feel alive.

The stress test: laughter mid-sentence
The laughing ad read
“No, stop, hahaha, okay, okay, I'm reading the ad now, I promise...”
First take, unedited, $0.016. The model did produce a laugh (a speech-to-text pass transcribes “Ha ha ha!”), and we can verify what it said but not how convincingly it said it. Play it and make the call: would you air this?

The pattern behind the seams: these models excel at prepared speech and struggle at the edges of performance. Real spontaneous laughter, crying, two voices talking over each other, the energy shift when a host is genuinely surprised: those come from a body and a moment, not from text. The model reads the sentence it’s given; it doesn’t know what’s at stake in the scene.

The subtler seam is fatigue. Ten seconds of synthetic speech is indistinguishable; the tenth hour is not, because human narrators drift and recover in ways that keep ears engaged, while a model’s habits repeat with perfect consistency. Listeners notice sameness before they can name it. For intros and explainers this barely matters. For a nine-hour novel it’s the whole question, and it is why we’d still audition any voice on a full chapter before committing a book to it.

Everything above used stock voices that ship with the model. The other path is cloning: give a model a sample of a specific person’s speech and it speaks as them. Cloning your own voice is a superpower, and the barrier is lower than most people think: instant cloning from about a minute of audio is included in ElevenLabs’ entry plans (paid tiers start around $5 a month, with professional-grade cloning from the $22 Creator tier), and newer models advertise working from even shorter samples. If you hate your recorded voice, or you need your podcast intro in a language you don’t speak, this is the feature you’re looking for.

The one piece of paper that makes someone else’s voice usable: a signed release. Illustration generated with Seedream 4.5 via Runware.

Anyone else’s voice is a different world, and the rule is short: written permission first, every time, no exceptions for parody or “it’s just for the group chat.” The law is moving one direction here. Tennessee’s ELVIS Act made unauthorized AI voice replicas actionable in March 2024 and other states have followed with their own digital-replica laws. The FCC ruled AI-generated voices in robocalls illegal under the TCPA in February 2024. The reason regulators moved fast is the dark side’s head start: the FTC has been warning since 2023 about scammers cloning a family member’s voice from a short social-media clip to stage fake emergencies. The same tool that narrates your audiobook can impersonate your kid. Treat it with the respect that implies.

Disclosure is the smaller cousin of consent, and platforms are formalizing it. YouTube requires creators to disclose realistic altered or synthetic content, and Audible labels every Virtual Voice title. Beyond the rules, it’s simply good practice: one caption line (“narration is AI-generated”) costs you nothing and protects the trust you’re building. This post generated its own examples and told you the model, the date, and the price of each one. That’s the habit.

What it costs, and a sensible first hour

The money question compresses to two routes. Metered APIs (the route we used) charge per character: about a dime a minute of finished audio at HD quality, with free-tier scraps on most platforms to experiment with. Subscription apps charge monthly: free tiers for tinkering, around $5 for entry plans with instant cloning, $22 for the tier serious creators actually use. Either way, the budget conversation for a normal project is cents and single dollars, not studios and day rates.

A first hour that teaches you the terrain:

  • Cast with a bake-off. Take one real paragraph you wrote and run it through three or four preset voices. Thirty cents, and you’ll immediately hear that voice choice matters more than model choice.
  • Direct one read. Respell a name, add an ellipsis, drop the speed to 0.95x. Compare takes. This is the actual skill, and it takes minutes to feel.
  • Ship something small. A podcast intro, a narrated product clip, one chapter of the family history project. Small enough to finish, real enough to judge.
  • If it should be your voice, clone yours. Record a clean minute, use a plan that supports it, and add the disclosure line. Never upload anyone else’s voice without their written permission.

The studio was never the point. The words were. For sixty years the gap between writing something and having it spoken well was money, gear, and other people’s time. As of about now, it’s punctuation.

Disclaimer: This article is general information, not legal advice, and reading it creates no attorney-client relationship. Laws, regulations, and court rulings summarized here reflect sources available as of August 2026 and may have changed. Consult counsel licensed in your jurisdiction before acting on any of it.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app