The best AI transcription models in 2026: one podcast, 16 models, one laptop
Every model aced the clean clip. Then two hosted Whisper copies quietly dropped whole sentences from the conversation.
On this page
He was a PhD physics, right, had a PhD in physics. His wife was a doctor, but he chose to work at inner-city school, and, you know, I think that was his public service.
He was, uh, he was a, a PhD physics, right? Had a PhD in physics. His wife was a doctor, but he chose to work at, uh, uh, inner city school, and, uh, you know, I think that was his public service.
He was a Ph.D. physics, right? Had a Ph.D. in physics. sentence dropped inner city school. And, you know, I think that was his public service.
He was a PhD physics, right? Had a PhD in physics. His wife was a doctor, but he chose to work at inner city school. And, you know, I think that was his public service.
We took five minutes of a real podcast and sent it to 16 cloud transcription models and one laptop. On the scripted intro, 437 words read into a good microphone, thirteen of the seventeen made one mistake or none. The whole experiment cost 21 cents. If your recordings sound like that intro, almost any model on this page will serve you, and the cheapest one costs about a cent per hour.
Most recordings do not sound like that. They are interviews, lectures and meetings, with two people talking over each other, half-finished sentences and names nobody spelled out. That is where the models separate, and not always in the order the price list suggests. For years the default answer to “what should I transcribe with” was Whisper. This post explains why that is no longer automatic, what to check instead, and which model fits which job.
On clean speech, every current model is close to perfect
The standard score for transcription is word error rate, or WER. Count every word the machine got wrong, left out or invented, and divide by the number of words actually spoken. A WER of 5% means one word in twenty is wrong. Anything under 5% reads as a finished transcript with a few typos.
Our test audio is episode 425 of NASA’s Houston We Have a Podcast, chosen because NASA material is in the public domain and the agency publishes a full transcript to score against. The first clip is the host’s scripted introduction: 138 seconds, one voice, studio sound, but dense with the things that trip transcribers. It has an episode number, “330 cubic feet,” an acronym (SCaN) and a guest’s name.
Six of seventeen runs matched the published transcript on every undisputed word. Seven more missed exactly one. The worst result was six errors in 437 words, a WER of 1.4%. Most of the misses were homophones: four models wrote “peaks inside the spacecraft” for “peeks,” and the two 0.6-billion-parameter models heard “lunar surface” as “learner surface.” Every run got the guest’s name and the acronym right.
| Model | Intro errors (437 words) | Conversation clip | Round trip | Billed per hour |
|---|---|---|---|---|
| MAI-Transcribe 2 | 0 | Verbatim: kept 48 “uh” and “um” | 6.0 s | $0.10 |
| Whisper Large V3, hosted | 0 | Dropped two passages, 23 words | 6.9 s | $0.06 |
| GPT-4o Transcribe | 0 | Clean, complete | 18.0 s | $0.22 |
| GPT Transcribe | 0 | Clean, complete | 15.2 s | $0.27 |
| Voxtral Small 24B | 0 | Clean, complete | 25.7 s | $0.18 |
| Whisper Turbo, local (whisper.cpp) | 0 | Clean, complete | 10.8 s | $0.00 |
| Qwen3 ASR 1.7B | 1 | Clean, complete | 6.9 s | $0.03 |
| Voxtral Mini Transcribe | 1 | Clean, complete | 7.9 s | $0.18 |
| Grok STT 1.0 | 1 | Clean, complete | 11.9 s | $0.10 |
| GPT-4o Mini Transcribe | 1 | Clean, complete | 11.6 s | $0.11 |
| Fish Audio Transcribe 1 | 1 | Clean, complete | 12.3 s | $0.36 |
| Deepgram Nova-3 | 1 | Complete; numbers spelled out as words | 5.5 s | $0.26 |
| Whisper Large V3 Turbo, hosted | 1 | Dropped three passages, 18 words | 7.8 s | $0.01 |
| Qwen3 ASR Flash | 2 | Clean, complete | 24.4 s | $0.13 |
| Whisper-1 | 2 | Added “Thank you” over the outro music | 19.1 s | $0.36 |
| Nemotron 3.5 ASR Streaming 0.6B | 4 | Wrote “Toule, Ohio” for Toledo | 12.8 s | $0.01 |
| Parakeet TDT 0.6B v3 | 6 | Verbatim: kept 51 “uh” and “um” | 5.2 s | $0.09 |
One caution about our own scoring. On one word, thirteen of sixteen cloud models agreed with each other and disagreed with NASA’s page. On another they split almost evenly. The published transcript has errors of its own, so we set both words aside. A small test like this can rank nothing past the first decimal place. What it can show is the shape of the field: on clean speech the gap between a free model and a 36-cent one is a homophone or two.
Conversation is where transcripts lose sentences
The second clip is 164 seconds of the host and his guest talking. It has interruptions, a guest who restarts sentences, and a hometown (Toledo, Ohio) that is easy to mishear. NASA’s transcript of this stretch has visible mistakes, so instead of a WER we checked specific things: the names, the places, and whether every spoken sentence made it onto the page.
Sixteen of seventeen runs wrote “Toledo.” One small streaming model wrote “Toule, Ohio.” The larger failure was quieter. The hosted copy of Whisper Large V3 Turbo dropped three passages, 18 words in total, including the whole clause “His wife was a doctor, but he chose to work at.” The hosted copy of the full Whisper Large V3 dropped two different passages, 23 words. Nothing in the output marks the gap. The sentence before runs straight into the sentence after, as the hero above shows. A reader would never know.
An error like that matters more than a misspelled word, because nobody proofreads for sentences that are not there. Whisper has a related, better-known habit: inventing text over silence or music. In our intro clip, OpenAI’s hosted Whisper-1 ended the transcript with a “Thank you” nobody said, and the Turbo copy appended the word “Music.” A 2024 study led by Allison Koenecke found that about 1% of Whisper transcriptions contained an entirely invented phrase or sentence, most often around long pauses. Newer models are built partly to fix this, and in our runs none of the non-Whisper models invented or dropped a sentence.
The other split is style. Microsoft’s MAI-Transcribe 2 and NVIDIA’s Parakeet kept every hesitation: 48 and 51 “uh” and “um” tokens in under three minutes. Every other model cleaned them out. Neither is wrong. A researcher coding an interview or a lawyer reviewing testimony wants verbatim. A podcaster writing show notes wants clean. Check which one a model gives you before you commit a backlog to it. Microsoft’s own API exposes a switch between “verbatim” and “clean”; through a router you get the default.
The leaderboards agree that Whisper is no longer on top
Five minutes of audio is a spot check. For ranking, use the public benchmarks, which test hours of speech across accents and recording conditions. They tell the same story from three directions.
Artificial Analysis scores commercial APIs on a blend of voice-agent calls, parliamentary speech and earnings calls. When we read it on September 21, the top five sat between 1.7% and 2.3% WER, with MAI-Transcribe 2 at 2.0% and ElevenLabs Scribe v2 at 2.2%. Whisper Large V3 scored 4.1%: still a usable transcript, but roughly twice the error rate of the leaders.
Among open models, Hugging Face’s Open ASR Leaderboard is the reference. NVIDIA’s Canary-Qwen 2.5B reports a mean WER of 5.63% there and Parakeet TDT 0.6B v3 6.34%, against 7.83% on the Whisper Large V3 Turbo model card. Those numbers come from harder test sets than Artificial Analysis uses, so compare within a leaderboard, never across two. We wrote more about that trap in how to read AI benchmarks.
OpenAI has moved on from Whisper too. Its pricing page now lists GPT Transcribe, a newer model that accepts context, keyword hints and several language hints at once, at a lower price than Whisper-1. Parakeet’s low rank in our small test next to its strong leaderboard score is a fair reminder of what a spot check is worth. It is a 0.6-billion-parameter model, and the errors it made were on exactly two phrases.
An hour of audio costs between one cent and 36 cents
Transcription is one of the few AI jobs priced in a unit everyone understands: the hour of audio. The spread is wide, and it does not track quality. The most expensive line on the table below is Whisper-1, the oldest model on it.
| Vendor | Model | 1 hour | 100 hours | Note |
|---|---|---|---|---|
| OpenAI | Whisper-1, GPT-4o Transcribe | $0.36 | $36.00 | $0.006 per minute; the diarizing GPT-4o tier costs the same |
| OpenAI | GPT Transcribe | $0.27 | $27.00 | $0.0045 per minute; GPT-4o Mini Transcribe is $0.18 per hour |
| Deepgram | Nova-3, pre-recorded | $0.26 | $25.80 | $0.0043 per minute pay as you go; $200 free credit |
| ElevenLabs | Scribe v2 | $0.22 | $22.00 | Speaker labels and word timestamps included |
| AssemblyAI | Universal-3.5 Pro | $0.21 | $21.00 | Speaker labels add $0.02 per hour; $50 free credit |
| Microsoft | MAI-Transcribe 2 | $0.10 | $10.00 | Launch price through December 31, 2026; no permanent price yet |
| Open weights, hosted | Whisper Large V3 Turbo | $0.01 | $1.20 | As billed through OpenRouter in our test |
| Open weights, local | Whisper Large V3 Turbo | $0.00 | $0.00 | A 547 MB download and your own electricity |
The sources are OpenAI, Deepgram, ElevenLabs, AssemblyAI and Microsoft. Two details deserve attention. Microsoft’s $0.10 per hour is a launch price that ends on December 31, 2026, and the company has not said what follows. And the one-cent rate for hosted Whisper Turbo is what OpenRouter’s transcription catalog billed us, not a promise: a router can send the same model name to different GPU hosts at different rates. Our two Whisper Large V3 calls were billed at rates four times apart.
For a person with a shoebox of recordings, the honest summary is that price barely matters. A 100-hour backlog, about two years of a weekly one-hour show, costs $36 at the top of the table and $10 to $27 for the models that lead the benchmarks. Price starts to matter at product scale, when the hours run into the thousands each month. Our cost of one hour of AI post puts these numbers next to text, image and video.
Speaker labels and timestamps decide more than accuracy does
When every model gets 98% of the words, the deciding question becomes what comes back besides the words. Three extras matter most.
Speaker labels. Called diarization, this marks who said what. For an interview it is the difference between a transcript and a wall of text. Scribe v2 labels up to 32 speakers at no extra charge. AssemblyAI adds it for $0.02 per hour. OpenAI sells it as a separate model, GPT-4o Transcribe Diarize, and MAI-Transcribe 2 added it at launch. Open-weight Whisper has none built in. One practical warning from our test: none of the 16 models returned speaker labels through the router’s endpoint. If you need them, call the vendor’s own API.
Timestamps. Subtitles and searchable audio need to know when each word was said. Scribe v2, MAI-Transcribe 2 and Parakeet return word-level times. Whisper’s standard output times each segment. OpenAI’s documentation for GPT Transcribe does not mention timestamps at all, so a subtitle workflow on OpenAI still runs through Whisper-1.
Languages. Whisper still has the widest reach, with 99 languages on its model card. Scribe v2 covers more than 90 and MAI-Transcribe 2 covers 60. Parakeet v3 handles 25 European languages, and Canary-Qwen, the open accuracy leader, is English only with a 40-second input limit. Outside English and the major European languages, test on your own audio before trusting any leaderboard. Length limits vary too: Scribe accepts files up to 3 GB and 10 hours, which saves you from cutting a long recording into pieces.
A laptop transcribes an hour of audio in about two minutes
Whisper may have lost the accuracy crown, but it is still the model you can download. The weights carry an MIT license, and whisper.cpp runs them on a Mac, a Windows PC or a Linux box with no Python and no account. We ran the Turbo model, compressed to a 547 MB file, on the same two clips.
The local run matched NASA’s transcript of the intro word for word, and it kept every passage that the hosted copies dropped. We cannot see inside the hosts, so we cannot say why. The lesson carries over either way: “Whisper” is not one product. The same weights behave differently depending on who runs them and how they split your audio.
At 28 times real time, an hour-long interview takes a little over two minutes on that laptop. The Whisper repository lists about 6 GB of memory for Turbo in its original form, and whisper.cpp lists under 1 GB for the 466 MB “small” model, which fits on almost any computer. Owners of an NVIDIA card have a second option in Parakeet v3, which is free for commercial use under CC BY 4.0, needs only 2 GB of memory and returns word timestamps.
The strongest argument for local is not the saved cent. It is that the recording never leaves the machine. A journalist’s source, a therapy session, a client call and a family argument all make poor uploads. CSuite’s audio workspace has a Transcribe mode built on this idea: Whisper Base, Small and Large V3 Turbo run on-device, and the same screen can send a file to OpenAI, ElevenLabs or OpenRouter when you want speaker labels or a cloud model’s accuracy. We listed other jobs that work with the network off in everyday wins for offline AI.
Pick by the job, then check the hard parts
Podcasters and interviewers need speaker labels first. Scribe v2 includes them, with word timestamps, at $0.22 per hour. AssemblyAI comes to $0.23 with labels added. Scribe v2 ranked fourth on Artificial Analysis when we checked. If the next step is a narrated version, our AI voiceover guide covers the other direction.
Journalists, researchers and anyone with sensitive audio should start local. Whisper Turbo through whisper.cpp was flawless on our clean clip and complete on the conversation. Add speaker names by hand, which for a two-person interview takes minutes.
Students with a lecture backlog can use the same local setup for free, or a hosted open model for about a cent per hour. Lectures are one voice into one microphone, the easy case.
Developers building a product should test MAI-Transcribe 2 and GPT Transcribe against their own audio, and budget for Microsoft’s price changing in January.
Whichever you pick, do not judge it on the first clean paragraph. Every model passes that test now. Take the messiest three minutes you have, the part with two people talking and a name nobody spelled, and read the output against the audio once. Look for the sentence that is missing, because that is the error no spell checker will ever find.


