The best AI voices for text to speech in 2026
Twelve voices read the same 70 words. Ten got every word, and an hour of speech costs anywhere from two cents to five dollars.
On this page
We wrote a 70-word podcast intro about weather forecasting and sent it to eleven text-to-speech models through the three platforms most people reach them on: Runware, Replicate and OpenRouter. Then we ran the same script through a free open-weight model on a laptop with the Wi-Fi off. All twelve came back with a human-sounding voice reading our words. Ten of them read every word. The whole experiment cost about 19 cents.
That is the good news, and it is also why picking a voice has become confusing. When every demo sounds fine, the decision moves to things a demo never shows: how fast the voice reads, what an hour of it costs, whether you can tell it how to say a line, how many languages it will say it in, and whether it works without a network. This post measures those five things on the same script. We do not rank naturalness for you. The clips are here; press play and rank them yourself.
Ten of twelve voices read the script word for word
The script is deliberately ordinary. It has a year (1950), a spelled-out number (twenty-four), a question, a list of three, and a quoted line. Those are the small traps that separate a model that reads from a model that recites. To check each read, we transcribed every clip back to text with Whisper on the same laptop and compared it with the script. Whisper is not a perfect judge, and the transcription post from two days ago shows the ways it slips, so we listened to the two clips it flagged and confirmed both by hand.
Ten reads matched the script on every word. The two exceptions were both open-weight models running on Replicate. Chatterbox read up to the quoted line and then stopped; the file has three seconds of silence where “Let’s start with the satellites” should be. Qwen3-TTS said “forecaster” as two words, which Whisper heard as “fear caster.” Both are fixable with a second attempt, but the rule of this test is one attempt, because that is what a script of a thousand lines gets.
Two of the cloud reads deserve a note. The MiniMax voice, whose preset is called Captivating Storyteller, was the slowest read at 129 words per minute and added a beat before the list. GPT Audio Mini was the fastest at 181 words per minute, and Whisper transcribed its question without a question mark: the voice ran “So what changed” straight into “three things.” Neither is an error. They are two different ideas of what a podcast intro should sound like, which is the point of listening before you buy.
Reading pace is the first difference you hear
Before you notice timbre or warmth, you notice speed. The twelve reads of the same 421 characters ran from 23.3 seconds to 32.5 seconds. That is a spread of 40% on identical text, and it is the difference between a voice that sounds like a news bulletin and one that sounds like a bedtime story. The hero chart at the top of this post is nothing more than those twelve clocks side by side.
Pace matters for money as well as mood. Every platform bills by characters or tokens of input, not by seconds of output, so a voice that reads slowly gives you more audio for the same bill. A ten-hour audiobook read at 129 words per minute is about 77,000 words; at 181 words per minute the same ten hours holds 109,000 words. Most models expose a speed parameter, and you should treat the default as a suggestion. Grok TTS, for instance, accepts a multiplier from 0.7 to 1.5 according to xAI’s documentation.
The other clock is the round trip: how long you wait for the file. It has nothing to do with pace and everything to do with where the model runs. ElevenLabs Flash v2.5 returned 30 seconds of audio in 2.4 seconds, in line with the roughly 75 milliseconds to first sound its models page promises. Gemini took 17.8 seconds and Qwen3-TTS 26.5 seconds for the same paragraph. For a batch job that renders a book overnight, none of that matters. For an assistant that has to answer while someone is listening, it is the whole product, and it is the one number in this post you should re-measure from your own region before you decide.
| Voice | Via | What Whisper heard | Words / min | Round trip | Per hour |
|---|---|---|---|---|---|
| ElevenLabs v3 | Replicate | Word for word | 138 | 8.7 s | $4.97 |
| GPT Audio | OpenRouter | Word for word | 138 | 6.3 s | $4.79 |
| MiniMax Speech 2.8 HD | Runware | Word for word | 129 | 9.6 s | $4.66 |
| ElevenLabs Flash v2.5 | Replicate | Word for word | 139 | 2.4 s | $2.51 |
| Gemini 3.1 Flash TTS | Runware | Word for word | 134 | 17.8 s | $1.81 |
| Chatterbox (Resemble AI) | Replicate | Stopped before the last sentence | 152 | 20.2 s | $1.37 |
| Qwen3-TTS 1.7B | Replicate | “Forecaster” came out as two words | 142 | 26.5 s | $1.02 |
| Grok TTS | Runware | Word for word | 152 | 5.9 s | $0.83 |
| Fish Audio S2.1 Pro | Runware | Word for word | 131 | 7.0 s | $0.71 |
| GPT Audio Mini | OpenRouter | Word for word, fastest read | 181 | 4.4 s | $0.23 |
| Kokoro 82M, hosted | Replicate | Word for word | 165 | 2.8 s | $0.02 |
| Kokoro 82M, local | This laptop | Word for word | 168 | 8.9 s | $0.00 |
An hour of speech costs between two cents and five dollars
The bill for an hour of finished audio is the number most buyers want and few vendors print, because every vendor bills a different unit. ElevenLabs, MiniMax, xAI and Fish Audio bill per character. OpenAI and Google bill per audio token, and Google’s pricing page states that a second of audio is 25 tokens. So we did the conversion the honest way: we took what each platform charged for our clip and scaled it to 60 minutes at that voice’s own pace.
At the top, ElevenLabs v3 at $0.10 per thousand characters on Replicate works out to $4.97 an hour, with GPT Audio at $4.79 and MiniMax Speech 2.8 HD at $4.66 close behind. MiniMax’s own price list is $100 per million characters for HD and $60 for Turbo, which matches what Runware billed us to the cent. In the middle, ElevenLabs Flash v2.5 is exactly half of v3, Gemini 3.1 Flash TTS lands at $1.81, and the hosted open-weight models cluster around a dollar.
Then there is the floor. Grok TTS and Fish Audio S2.1 Pro both list at $15 per million characters, about 80 cents an hour. GPT Audio Mini, billed at $2.40 per million audio tokens on OpenRouter, cost us 23 cents an hour. Kokoro on Replicate cost two cents an hour because it bills for 0.8 seconds of a shared GPU. The gap between the cheapest and the priciest full-service cloud voice is more than 20 times, a hosted open model sits another ten times below that, and the gap to a local model is infinite in the arithmetic sense. Our cost-of-one-hour post found the same shape in text and images: the price of the premium option is set by what people will pay, not by what it costs to run.
One trap in the arithmetic: ElevenLabs sells subscriptions, not characters. Its pricing page lists a $22 Creator plan with 121,000 credits, one credit per character on v3, which it describes as about 121 minutes of speech. That is close to $11 an hour if you use every credit and more if you do not, so the Replicate rate is the cheaper way to buy the same model in small amounts. The audiobook narration guide covers the licensing side of that choice, which matters more than the price for anything you sell.
Tags, instructions and cloning matter more than the preset voice
A preset voice is the least interesting thing about a speech model, because every vendor now ships dozens. MiniMax’s API schema on Runware lists 193 of them, Google’s lists 30, xAI’s lists 27. What separates the models is how you steer a line once you have picked a voice, and there are three approaches.
The first is inline tags. ElevenLabs v3 reads audio tags such as [whispers] and [sighs] inside the text, per its models page. Fish Audio takes free-form bracket cues like [laughing nervously] and multi-speaker tags in one request, per its model overview. Gemini’s schema accepts [Sam] and [Bob] speaker labels for dialogue. The second approach is a separate instruction. OpenAI’s GPT-4o Mini TTS takes a free-text description of tone, accent and pacing alongside the script, and Gemini takes a style prompt the same way. The third is cloning: Fish Audio, MiniMax, Qwen3-TTS and Chatterbox all build a voice from a short recording, and Qwen’s model card claims it needs three seconds of it.
| Model | Languages | Steering | Voice cloning | Per request |
|---|---|---|---|---|
| ElevenLabs v3 | 70+ | Audio tags such as [whispers], [sighs]; stability slider | Yes | 5,000 chars |
| GPT-4o Mini TTS | Not stated | Free-text instructions for tone, accent, pacing | No | 2,000 tokens |
| MiniMax Speech 2.8 | 40+ boost codes | Emotion presets, pause markers, 193 voices | $1.50 per voice | 50,000 chars |
| Gemini 3.1 Flash TTS | 70+ codes | Style prompt in plain words; [Sam] [Bob] dialogue | No | Not stated |
| Grok TTS | 20 | Speed 0.7 to 1.5; inline tags such as [pause] | No | 60,000 chars |
| Fish Audio S2.1 Pro | 83 | Free-form [bracket] cues; multi-speaker tags | Zero-shot, from a reference clip | Not stated |
| Qwen3-TTS (open) | 10 | Style instruction; voice design from a description | 3-second clone | Your hardware |
| Chatterbox (open) | 23 | Exaggeration and pace weights | From a short clip | Your hardware |
| Kokoro 82M (open) | 8 | Speed only; 54 preset voices | No | Your hardware |
Which approach you need depends on the job. A narrator reading a novel wants tags, because emotion changes line by line. A product that speaks to customers wants an instruction, because the whole voice should sound one way. A creator who is the voice of their own channel wants cloning, and should read the cloning rules before recording anything, because a cloned voice you do not own is a legal problem rather than a technical one. Request limits are the quiet fourth column: ElevenLabs v3 stops at 5,000 characters per call, MiniMax at 50,000, xAI at 60,000. For a chapter, that decides how many stitches your pipeline needs.
Language counts run from 8 to 83, and the count is not the quality
Fish Audio says 83 languages. ElevenLabs v3 says 70 or more. Chatterbox Multilingual lists 23, xAI lists 20, Qwen3-TTS lists 10, and Kokoro’s model card lists 8. Those numbers are true and they are also the least useful thing on a spec sheet, for the same reason a phrasebook is not fluency. A model that lists Hindi may read Hindi with an accent that a Hindi speaker finds unusable, and no vendor publishes a per-language quality score.
The practical test is the one we ran here, in the language you need: one paragraph with a number, a name and a question, sent to the three or four candidates, judged by a native speaker. Fish Audio and xAI both make that cheap, at less than a cent per paragraph. Two other details matter more than the count. Whether the model detects the language itself (xAI, MiniMax and Qwen all accept an “auto” setting) decides how mixed-language text behaves. And whether the language list includes the specific accent, such as Portuguese from Brazil versus Portugal, which xAI’s list keeps separate.
A laptop reads the paragraph in nine seconds with no account
The cheapest voice in this test is free, and it is not a toy. Kokoro is an 82-million-parameter model released under Apache 2.0, with 54 preset voices. We ran it on a MacBook Pro through the kokoro-js package, which loads quantized ONNX weights and runs them on the CPU with no GPU involved. The 70-word script took 8.9 seconds to generate, about three times faster than real time. Whisper transcribed the result word for word, and its length matched the hosted copy of the same model on Replicate to within half a second.
Three other open models are worth knowing. Qwen3-TTS, from Alibaba, ships 0.6 and 1.7-billion-parameter versions under Apache 2.0 and adds cloning and voice design from a text description; it was the slowest hosted read in our test at 26.5 seconds of round trip but the most capable open model on the control table. Chatterbox is MIT-licensed, 500 million parameters, and stamps every file with an inaudible watermark, which is a feature or a constraint depending on your use. And Supertonic, a 66-million-parameter model that its authors measured at 912 characters per second on an M4 Pro CPU, is the one CSuite bundles for on-device speech, because it runs inside the app with no key at all.
The trade is control and reach for privacy and price. None of the three local models steer with tags the way ElevenLabs or Fish do, and the local language lists are short. But a script that never leaves the machine is the only kind that a confidentiality clause allows, and a voice that works on a train is the only kind that still works when the network does not.
Pick by the job, then listen
If the voice is the product, as in an audiobook or a branded show, the top tier earns its price: ElevenLabs v3 for tag-level control, GPT Audio or GPT-4o Mini TTS for one instruction that shapes everything, MiniMax for the largest voice menu and long requests. Budget about $5 an hour of finished audio and read the license before you sell it.
If the voice is a feature, as in an app that reads notifications, a course that needs narration in six languages, or a video that needs a scratch track, start at the floor. Fish Audio S2.1 Pro and Grok TTS cost under a dollar an hour, read our script cleanly, and steer well enough for most lines. GPT Audio Mini is cheaper still if you can live with its pace or slow it down.
If the text is private, or the machine is offline, or the budget is zero, run Kokoro. It read our script without an error in nine seconds on a laptop CPU, and the voiceover workflow post shows how to turn that into a finished track. Whatever tier you land on, do what this post did before committing a project to a voice: one paragraph with a number, a question and a quote, one take, and your own ears.


