Skip to content
CSuite
Sep 23, 202610 min read

The best AI voices for text to speech in 2026

Twelve voices read the same 70 words. Ten got every word, and an hour of speech costs anywhere from two cents to five dollars.

On this page
One 70-word script · 11 cloud voices · one laptop · September 2026
Every voice read the same paragraph. The clock and the bill are where they differ.
GPT Audio Mini
23.3 s
$0.23
Kokoro 82M, on this laptop
25.0 s
$0
Kokoro 82M, hosted
25.5 s
$0.02
Grok TTS
27.6 s
$0.83
Chatterbox
27.6 s
$1.37
Qwen3-TTS 1.7B
29.6 s
$1.02
ElevenLabs Flash v2.5
30.1 s
$2.51
GPT Audio
30.5 s
$4.79
ElevenLabs v3
30.5 s
$4.97
Gemini 3.1 Flash TTS
31.3 s
$1.81
Fish Audio S2.1 Pro
32.0 s
$0.71
MiniMax Speech 2.8 HD
32.5 s
$4.66
The bar is how long each voice took to read the same 421-character script, from 23 to 33 seconds. The amber figure is what an hour of speech would cost at that pace and at the rate we were billed, or the list rate where the platform bills per character. Generated on September 23, 2026 through Runware, Replicate and OpenRouter, default voices, first result, no rerolls. The whole experiment cost about 19 cents.

We wrote a 70-word podcast intro about weather forecasting and sent it to eleven text-to-speech models through the three platforms most people reach them on: Runware, Replicate and OpenRouter. Then we ran the same script through a free open-weight model on a laptop with the Wi-Fi off. All twelve came back with a human-sounding voice reading our words. Ten of them read every word. The whole experiment cost about 19 cents.

That is the good news, and it is also why picking a voice has become confusing. When every demo sounds fine, the decision moves to things a demo never shows: how fast the voice reads, what an hour of it costs, whether you can tell it how to say a line, how many languages it will say it in, and whether it works without a network. This post measures those five things on the same script. We do not rank naturalness for you. The clips are here; press play and rank them yourself.

Ten of twelve voices read the script word for word

The script is deliberately ordinary. It has a year (1950), a spelled-out number (twenty-four), a question, a list of three, and a quoted line. Those are the small traps that separate a model that reads from a model that recites. To check each read, we transcribed every clip back to text with Whisper on the same laptop and compared it with the script. Whisper is not a perfect judge, and the transcription post from two days ago shows the ways it slips, so we listened to the two clips it flagged and confirmed both by hand.

Ten reads matched the script on every word. The two exceptions were both open-weight models running on Replicate. Chatterbox read up to the quoted line and then stopped; the file has three seconds of silence where “Let’s start with the satellites” should be. Qwen3-TTS said “forecaster” as two words, which Whisper heard as “fear caster.” Both are fixable with a second attempt, but the rule of this test is one attempt, because that is what a script of a thousand lines gets.

The same script, eight cloud voices · one take each
ElevenLabs v3 (Rachel)
$0.042 at Replicate’s list rate
0:00 / 0:31
GPT Audio (alloy)
$0.041 billed via OpenRouter
0:00 / 0:31
MiniMax Speech 2.8 HD (Captivating Storyteller)
$0.042 billed via Runware
0:00 / 0:33
ElevenLabs Flash v2.5 (Rachel)
$0.021 at Replicate’s list rate
0:00 / 0:30
Gemini 3.1 Flash TTS (Kore)
$0.016 billed via Runware
0:00 / 0:31
Grok TTS (eve)
$0.006 billed via Runware
0:00 / 0:28
Fish Audio S2.1 Pro (Adrian)
$0.006 billed via Runware
0:00 / 0:32
GPT Audio Mini (alloy)
$0.0015 billed via OpenRouter
0:00 / 0:23
Generated on September 23, 2026 with each platform’s default voice and settings, one request per model, no rerolls, no editing beyond re-encoding to MP3. Listen for the pause after “So what changed?”, the list of three, and whether the quoted line sounds quoted.

Two of the cloud reads deserve a note. The MiniMax voice, whose preset is called Captivating Storyteller, was the slowest read at 129 words per minute and added a beat before the list. GPT Audio Mini was the fastest at 181 words per minute, and Whisper transcribed its question without a question mark: the voice ran “So what changed” straight into “three things.” Neither is an error. They are two different ideas of what a podcast intro should sound like, which is the point of listening before you buy.

One script on the stand, many voices on the timeline: the fair way to compare is to change nothing but the model. Illustration generated with Seedream 5 Pro via Runware.

Reading pace is the first difference you hear

Before you notice timbre or warmth, you notice speed. The twelve reads of the same 421 characters ran from 23.3 seconds to 32.5 seconds. That is a spread of 40% on identical text, and it is the difference between a voice that sounds like a news bulletin and one that sounds like a bedtime story. The hero chart at the top of this post is nothing more than those twelve clocks side by side.

Pace matters for money as well as mood. Every platform bills by characters or tokens of input, not by seconds of output, so a voice that reads slowly gives you more audio for the same bill. A ten-hour audiobook read at 129 words per minute is about 77,000 words; at 181 words per minute the same ten hours holds 109,000 words. Most models expose a speed parameter, and you should treat the default as a suggestion. Grok TTS, for instance, accepts a multiplier from 0.7 to 1.5 according to xAI’s documentation.

The other clock is the round trip: how long you wait for the file. It has nothing to do with pace and everything to do with where the model runs. ElevenLabs Flash v2.5 returned 30 seconds of audio in 2.4 seconds, in line with the roughly 75 milliseconds to first sound its models page promises. Gemini took 17.8 seconds and Qwen3-TTS 26.5 seconds for the same paragraph. For a batch job that renders a book overnight, none of that matters. For an assistant that has to answer while someone is listening, it is the whole product, and it is the one number in this post you should re-measure from your own region before you decide.

Same 70 words, twelve reads
VoiceViaWhat Whisper heardWords / minRound tripPer hour
ElevenLabs v3ReplicateWord for word1388.7 s$4.97
GPT AudioOpenRouterWord for word1386.3 s$4.79
MiniMax Speech 2.8 HDRunwareWord for word1299.6 s$4.66
ElevenLabs Flash v2.5ReplicateWord for word1392.4 s$2.51
Gemini 3.1 Flash TTSRunwareWord for word13417.8 s$1.81
Chatterbox (Resemble AI)ReplicateStopped before the last sentence15220.2 s$1.37
Qwen3-TTS 1.7BReplicate“Forecaster” came out as two words14226.5 s$1.02
Grok TTSRunwareWord for word1525.9 s$0.83
Fish Audio S2.1 ProRunwareWord for word1317.0 s$0.71
GPT Audio MiniOpenRouterWord for word, fastest read1814.4 s$0.23
Kokoro 82M, hostedReplicateWord for word1652.8 s$0.02
Kokoro 82M, localThis laptopWord for word1688.9 s$0.00
Each clip was transcribed back with Whisper Large V3 Turbo on the same laptop and compared with the script. Whisper writes “24” for “twenty-four” and adds a stray “Thank you” over trailing silence; neither is counted. Words per minute is 70 words divided by the clip length. Round trip is request to file from one desk, not a latency benchmark. Per hour scales what we were billed (Runware and OpenRouter report the charge; Replicate bills per character at its list rate) to 60 minutes at each voice’s own pace. Run on September 23, 2026, first result shown.

An hour of speech costs between two cents and five dollars

The bill for an hour of finished audio is the number most buyers want and few vendors print, because every vendor bills a different unit. ElevenLabs, MiniMax, xAI and Fish Audio bill per character. OpenAI and Google bill per audio token, and Google’s pricing page states that a second of audio is 25 tokens. So we did the conversion the honest way: we took what each platform charged for our clip and scaled it to 60 minutes at that voice’s own pace.

At the top, ElevenLabs v3 at $0.10 per thousand characters on Replicate works out to $4.97 an hour, with GPT Audio at $4.79 and MiniMax Speech 2.8 HD at $4.66 close behind. MiniMax’s own price list is $100 per million characters for HD and $60 for Turbo, which matches what Runware billed us to the cent. In the middle, ElevenLabs Flash v2.5 is exactly half of v3, Gemini 3.1 Flash TTS lands at $1.81, and the hosted open-weight models cluster around a dollar.

Then there is the floor. Grok TTS and Fish Audio S2.1 Pro both list at $15 per million characters, about 80 cents an hour. GPT Audio Mini, billed at $2.40 per million audio tokens on OpenRouter, cost us 23 cents an hour. Kokoro on Replicate cost two cents an hour because it bills for 0.8 seconds of a shared GPU. The gap between the cheapest and the priciest full-service cloud voice is more than 20 times, a hosted open model sits another ten times below that, and the gap to a local model is infinite in the arithmetic sense. Our cost-of-one-hour post found the same shape in text and images: the price of the premium option is set by what people will pay, not by what it costs to run.

One trap in the arithmetic: ElevenLabs sells subscriptions, not characters. Its pricing page lists a $22 Creator plan with 121,000 credits, one credit per character on v3, which it describes as about 121 minutes of speech. That is close to $11 an hour if you use every credit and more if you do not, so the Replicate rate is the cheaper way to buy the same model in small amounts. The audiobook narration guide covers the licensing side of that choice, which matters more than the price for anything you sell.

Sixty minutes of finished audio is the unit that matters, and the receipt for it runs from two cents to five dollars. Illustration generated with Seedream 5 Pro via Runware.

Tags, instructions and cloning matter more than the preset voice

A preset voice is the least interesting thing about a speech model, because every vendor now ships dozens. MiniMax’s API schema on Runware lists 193 of them, Google’s lists 30, xAI’s lists 27. What separates the models is how you steer a line once you have picked a voice, and there are three approaches.

The first is inline tags. ElevenLabs v3 reads audio tags such as [whispers] and [sighs] inside the text, per its models page. Fish Audio takes free-form bracket cues like [laughing nervously] and multi-speaker tags in one request, per its model overview. Gemini’s schema accepts [Sam] and [Bob] speaker labels for dialogue. The second approach is a separate instruction. OpenAI’s GPT-4o Mini TTS takes a free-text description of tone, accent and pacing alongside the script, and Gemini takes a style prompt the same way. The third is cloning: Fish Audio, MiniMax, Qwen3-TTS and Chatterbox all build a voice from a short recording, and Qwen’s model card claims it needs three seconds of it.

What each voice lets you change
ModelLanguagesSteeringVoice cloningPer request
ElevenLabs v370+Audio tags such as [whispers], [sighs]; stability sliderYes5,000 chars
GPT-4o Mini TTSNot statedFree-text instructions for tone, accent, pacingNo2,000 tokens
MiniMax Speech 2.840+ boost codesEmotion presets, pause markers, 193 voices$1.50 per voice50,000 chars
Gemini 3.1 Flash TTS70+ codesStyle prompt in plain words; [Sam] [Bob] dialogueNoNot stated
Grok TTS20Speed 0.7 to 1.5; inline tags such as [pause]No60,000 chars
Fish Audio S2.1 Pro83Free-form [bracket] cues; multi-speaker tagsZero-shot, from a reference clipNot stated
Qwen3-TTS (open)10Style instruction; voice design from a description3-second cloneYour hardware
Chatterbox (open)23Exaggeration and pace weightsFrom a short clipYour hardware
Kokoro 82M (open)8Speed only; 54 preset voicesNoYour hardware
Read from each vendor’s documentation, model card or API schema on September 23, 2026. Language counts are the vendor’s own claim, and a listed language is not a promise of a good accent in it.

Which approach you need depends on the job. A narrator reading a novel wants tags, because emotion changes line by line. A product that speaks to customers wants an instruction, because the whole voice should sound one way. A creator who is the voice of their own channel wants cloning, and should read the cloning rules before recording anything, because a cloned voice you do not own is a legal problem rather than a technical one. Request limits are the quiet fourth column: ElevenLabs v3 stops at 5,000 characters per call, MiniMax at 50,000, xAI at 60,000. For a chapter, that decides how many stitches your pipeline needs.

Language counts run from 8 to 83, and the count is not the quality

Fish Audio says 83 languages. ElevenLabs v3 says 70 or more. Chatterbox Multilingual lists 23, xAI lists 20, Qwen3-TTS lists 10, and Kokoro’s model card lists 8. Those numbers are true and they are also the least useful thing on a spec sheet, for the same reason a phrasebook is not fluency. A model that lists Hindi may read Hindi with an accent that a Hindi speaker finds unusable, and no vendor publishes a per-language quality score.

The practical test is the one we ran here, in the language you need: one paragraph with a number, a name and a question, sent to the three or four candidates, judged by a native speaker. Fish Audio and xAI both make that cheap, at less than a cent per paragraph. Two other details matter more than the count. Whether the model detects the language itself (xAI, MiniMax and Qwen all accept an “auto” setting) decides how mixed-language text behaves. And whether the language list includes the specific accent, such as Portuguese from Brazil versus Portugal, which xAI’s list keeps separate.

A long language list is a menu, not a guarantee. The paragraph test in the language you need is the only spec that counts. Illustration generated with Seedream 5 Pro via Runware.

A laptop reads the paragraph in nine seconds with no account

The cheapest voice in this test is free, and it is not a toy. Kokoro is an 82-million-parameter model released under Apache 2.0, with 54 preset voices. We ran it on a MacBook Pro through the kokoro-js package, which loads quantized ONNX weights and runs them on the CPU with no GPU involved. The 70-word script took 8.9 seconds to generate, about three times faster than real time. Whisper transcribed the result word for word, and its length matched the hosted copy of the same model on Replicate to within half a second.

Two commands, no account, no upload
$ npm install kokoro-js
$ node -e "import('kokoro-js').then(async ({KokoroTTS}) => { const tts = await KokoroTTS.from_pretrained('onnx-community/Kokoro-82M-v1.0-ONNX', {dtype:'q8'}); const audio = await tts.generate('Welcome back to the show.', {voice:'af_bella'}); await audio.save('read.wav'); })"
Run as written on a MacBook Pro with an M3 Max and 36 GB of memory, Node 26. The first run downloads about 90 MB of quantized weights and took 85 seconds to load; every run after that loaded in 0.3 seconds. Reading the full 70-word script took 8.9 seconds on the CPU alone and produced 25 seconds of audio.

Three other open models are worth knowing. Qwen3-TTS, from Alibaba, ships 0.6 and 1.7-billion-parameter versions under Apache 2.0 and adds cloning and voice design from a text description; it was the slowest hosted read in our test at 26.5 seconds of round trip but the most capable open model on the control table. Chatterbox is MIT-licensed, 500 million parameters, and stamps every file with an inaudible watermark, which is a feature or a constraint depending on your use. And Supertonic, a 66-million-parameter model that its authors measured at 912 characters per second on an M4 Pro CPU, is the one CSuite bundles for on-device speech, because it runs inside the app with no key at all.

The trade is control and reach for privacy and price. None of the three local models steer with tags the way ElevenLabs or Fish do, and the local language lists are short. But a script that never leaves the machine is the only kind that a confidentiality clause allows, and a voice that works on a train is the only kind that still works when the network does not.

The same script, four open-weight reads · one take each
Kokoro 82M (af_bella), on this laptop’s CPU
8.9 s to generate, $0
0:00 / 0:25
Kokoro 82M (af_bella), hosted on Replicate
0.8 s of GPU time, about $0.0002
0:00 / 0:25
Qwen3-TTS 1.7B (Serena), hosted on Replicate
$0.008 at list rate
0:00 / 0:30
Chatterbox (default voice), hosted on Replicate
$0.011 at list rate; last sentence missing
0:00 / 0:28
Same script and date. The local Kokoro clip was made with the kokoro-js package running the quantized ONNX weights on an Apple M3 Max CPU, no GPU. The hosted Kokoro clip is the same weights on Replicate; the two reads are the same length to within half a second.
Nine seconds of CPU time for a paragraph, no network, no account: the open-weight voice is now inside the pack, not behind it. Illustration generated with Seedream 5 Pro via Runware.

Pick by the job, then listen

If the voice is the product, as in an audiobook or a branded show, the top tier earns its price: ElevenLabs v3 for tag-level control, GPT Audio or GPT-4o Mini TTS for one instruction that shapes everything, MiniMax for the largest voice menu and long requests. Budget about $5 an hour of finished audio and read the license before you sell it.

If the voice is a feature, as in an app that reads notifications, a course that needs narration in six languages, or a video that needs a scratch track, start at the floor. Fish Audio S2.1 Pro and Grok TTS cost under a dollar an hour, read our script cleanly, and steer well enough for most lines. GPT Audio Mini is cheaper still if you can live with its pace or slow it down.

If the text is private, or the machine is offline, or the budget is zero, run Kokoro. It read our script without an error in nine seconds on a laptop CPU, and the voiceover workflow post shows how to turn that into a finished track. Whatever tier you land on, do what this post did before committing a project to a voice: one paragraph with a number, a question and a quote, one take, and your own ears.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app