Skip to content
CSuite
Sep 21, 202610 min read

The best AI transcription models in 2026: one podcast, 16 models, one laptop

Every model aced the clean clip. Then two hosted Whisper copies quietly dropped whole sentences from the conversation.

On this page
One podcast · 16 cloud models · one laptop · September 2026
Every model heard the scripted intro. The conversation is where they split.
GPT Transcribe
Clean read, everything kept

He was a PhD physics, right, had a PhD in physics. His wife was a doctor, but he chose to work at inner-city school, and, you know, I think that was his public service.

MAI-Transcribe 2
Verbatim, every filler kept

He was, uh, he was a, a PhD physics, right? Had a PhD in physics. His wife was a doctor, but he chose to work at, uh, uh, inner city school, and, uh, you know, I think that was his public service.

Whisper Large V3 Turbo, hosted
Eleven words missing

He was a Ph.D. physics, right? Had a Ph.D. in physics. sentence dropped inner city school. And, you know, I think that was his public service.

Whisper Large V3 Turbo, on a laptop
Same weights, nothing missing

He was a PhD physics, right? Had a PhD in physics. His wife was a doctor, but he chose to work at inner city school. And, you know, I think that was his public service.

The same 15 seconds of NASA’s public-domain podcast, transcribed on September 21, 2026. Cloud runs went through OpenRouter’s transcription endpoint with default settings, one attempt per model, no retries. All 32 cloud runs cost $0.21 in total.

We took five minutes of a real podcast and sent it to 16 cloud transcription models and one laptop. On the scripted intro, 437 words read into a good microphone, thirteen of the seventeen made one mistake or none. The whole experiment cost 21 cents. If your recordings sound like that intro, almost any model on this page will serve you, and the cheapest one costs about a cent per hour.

Most recordings do not sound like that. They are interviews, lectures and meetings, with two people talking over each other, half-finished sentences and names nobody spelled out. That is where the models separate, and not always in the order the price list suggests. For years the default answer to “what should I transcribe with” was Whisper. This post explains why that is no longer automatic, what to check instead, and which model fits which job.

On clean speech, every current model is close to perfect

The standard score for transcription is word error rate, or WER. Count every word the machine got wrong, left out or invented, and divide by the number of words actually spoken. A WER of 5% means one word in twenty is wrong. Anything under 5% reads as a finished transcript with a few typos.

Our test audio is episode 425 of NASA’s Houston We Have a Podcast, chosen because NASA material is in the public domain and the agency publishes a full transcript to score against. The first clip is the host’s scripted introduction: 138 seconds, one voice, studio sound, but dense with the things that trip transcribers. It has an episode number, “330 cubic feet,” an acronym (SCaN) and a guest’s name.

Six of seventeen runs matched the published transcript on every undisputed word. Seven more missed exactly one. The worst result was six errors in 437 words, a WER of 1.4%. Most of the misses were homophones: four models wrote “peaks inside the spacecraft” for “peeks,” and the two 0.6-billion-parameter models heard “lunar surface” as “learner surface.” Every run got the guest’s name and the acronym right.

Same five minutes of audio, 17 transcribers
ModelIntro errors (437 words)Conversation clipRound tripBilled per hour
MAI-Transcribe 20Verbatim: kept 48 “uh” and “um”6.0 s$0.10
Whisper Large V3, hosted0Dropped two passages, 23 words6.9 s$0.06
GPT-4o Transcribe0Clean, complete18.0 s$0.22
GPT Transcribe0Clean, complete15.2 s$0.27
Voxtral Small 24B0Clean, complete25.7 s$0.18
Whisper Turbo, local (whisper.cpp)0Clean, complete10.8 s$0.00
Qwen3 ASR 1.7B1Clean, complete6.9 s$0.03
Voxtral Mini Transcribe1Clean, complete7.9 s$0.18
Grok STT 1.01Clean, complete11.9 s$0.10
GPT-4o Mini Transcribe1Clean, complete11.6 s$0.11
Fish Audio Transcribe 11Clean, complete12.3 s$0.36
Deepgram Nova-31Complete; numbers spelled out as words5.5 s$0.26
Whisper Large V3 Turbo, hosted1Dropped three passages, 18 words7.8 s$0.01
Qwen3 ASR Flash2Clean, complete24.4 s$0.13
Whisper-12Added “Thank you” over the outro music19.1 s$0.36
Nemotron 3.5 ASR Streaming 0.6B4Wrote “Toule, Ohio” for Toledo12.8 s$0.01
Parakeet TDT 0.6B v36Verbatim: kept 51 “uh” and “um”5.2 s$0.09
Two clips from NASA’s Houston We Have a Podcast, episode 425: a 138-second scripted intro and a 164-second two-person conversation. Errors are counted against NASA’s published transcript after lowercasing and stripping punctuation. Two words where the models split against the published transcript (“his” or “its”, “help” or “helped”) are excluded as disputed, and spelled-out numbers are not counted as errors. Round trip is upload plus processing for both clips from one desk, not a speed benchmark. Billed per hour is the actual OpenRouter charge scaled to 60 minutes. Run on September 21, 2026, first result shown.

One caution about our own scoring. On one word, thirteen of sixteen cloud models agreed with each other and disagreed with NASA’s page. On another they split almost evenly. The published transcript has errors of its own, so we set both words aside. A small test like this can rank nothing past the first decimal place. What it can show is the shape of the field: on clean speech the gap between a free model and a 36-cent one is a homophone or two.

One voice, one good microphone: the kind of recording every current model turns into a near-finished transcript. Illustration generated with Seedream 5 Pro via Runware.

Conversation is where transcripts lose sentences

The second clip is 164 seconds of the host and his guest talking. It has interruptions, a guest who restarts sentences, and a hometown (Toledo, Ohio) that is easy to mishear. NASA’s transcript of this stretch has visible mistakes, so instead of a WER we checked specific things: the names, the places, and whether every spoken sentence made it onto the page.

Sixteen of seventeen runs wrote “Toledo.” One small streaming model wrote “Toule, Ohio.” The larger failure was quieter. The hosted copy of Whisper Large V3 Turbo dropped three passages, 18 words in total, including the whole clause “His wife was a doctor, but he chose to work at.” The hosted copy of the full Whisper Large V3 dropped two different passages, 23 words. Nothing in the output marks the gap. The sentence before runs straight into the sentence after, as the hero above shows. A reader would never know.

An error like that matters more than a misspelled word, because nobody proofreads for sentences that are not there. Whisper has a related, better-known habit: inventing text over silence or music. In our intro clip, OpenAI’s hosted Whisper-1 ended the transcript with a “Thank you” nobody said, and the Turbo copy appended the word “Music.” A 2024 study led by Allison Koenecke found that about 1% of Whisper transcriptions contained an entirely invented phrase or sentence, most often around long pauses. Newer models are built partly to fix this, and in our runs none of the non-Whisper models invented or dropped a sentence.

The other split is style. Microsoft’s MAI-Transcribe 2 and NVIDIA’s Parakeet kept every hesitation: 48 and 51 “uh” and “um” tokens in under three minutes. Every other model cleaned them out. Neither is wrong. A researcher coding an interview or a lawyer reviewing testimony wants verbatim. A podcaster writing show notes wants clean. Check which one a model gives you before you commit a backlog to it. Microsoft’s own API exposes a switch between “verbatim” and “clean”; through a router you get the default.

Two people, restarts and crosstalk: the recording that separates models the clean clip could not. Illustration generated with Seedream 5 Pro via Runware.

The leaderboards agree that Whisper is no longer on top

Five minutes of audio is a spot check. For ranking, use the public benchmarks, which test hours of speech across accents and recording conditions. They tell the same story from three directions.

Artificial Analysis scores commercial APIs on a blend of voice-agent calls, parliamentary speech and earnings calls. When we read it on September 21, the top five sat between 1.7% and 2.3% WER, with MAI-Transcribe 2 at 2.0% and ElevenLabs Scribe v2 at 2.2%. Whisper Large V3 scored 4.1%: still a usable transcript, but roughly twice the error rate of the leaders.

Among open models, Hugging Face’s Open ASR Leaderboard is the reference. NVIDIA’s Canary-Qwen 2.5B reports a mean WER of 5.63% there and Parakeet TDT 0.6B v3 6.34%, against 7.83% on the Whisper Large V3 Turbo model card. Those numbers come from harder test sets than Artificial Analysis uses, so compare within a leaderboard, never across two. We wrote more about that trap in how to read AI benchmarks.

OpenAI has moved on from Whisper too. Its pricing page now lists GPT Transcribe, a newer model that accepts context, keyword hints and several language hints at once, at a lower price than Whisper-1. Parakeet’s low rank in our small test next to its strong leaderboard score is a fair reminder of what a spot check is worth. It is a 0.6-billion-parameter model, and the errors it made were on exactly two phrases.

An hour of audio costs between one cent and 36 cents

Transcription is one of the few AI jobs priced in a unit everyone understands: the hour of audio. The spread is wide, and it does not track quality. The most expensive line on the table below is Whisper-1, the oldest model on it.

What one hour, and a 100-hour backlog, costs
VendorModel1 hour100 hoursNote
OpenAIWhisper-1, GPT-4o Transcribe$0.36$36.00$0.006 per minute; the diarizing GPT-4o tier costs the same
OpenAIGPT Transcribe$0.27$27.00$0.0045 per minute; GPT-4o Mini Transcribe is $0.18 per hour
DeepgramNova-3, pre-recorded$0.26$25.80$0.0043 per minute pay as you go; $200 free credit
ElevenLabsScribe v2$0.22$22.00Speaker labels and word timestamps included
AssemblyAIUniversal-3.5 Pro$0.21$21.00Speaker labels add $0.02 per hour; $50 free credit
MicrosoftMAI-Transcribe 2$0.10$10.00Launch price through December 31, 2026; no permanent price yet
Open weights, hostedWhisper Large V3 Turbo$0.01$1.20As billed through OpenRouter in our test
Open weights, localWhisper Large V3 Turbo$0.00$0.00A 547 MB download and your own electricity
List prices read from each vendor’s pricing page on September 21, 2026, for pre-recorded files. Live streaming tiers cost more everywhere.

The sources are OpenAI, Deepgram, ElevenLabs, AssemblyAI and Microsoft. Two details deserve attention. Microsoft’s $0.10 per hour is a launch price that ends on December 31, 2026, and the company has not said what follows. And the one-cent rate for hosted Whisper Turbo is what OpenRouter’s transcription catalog billed us, not a promise: a router can send the same model name to different GPU hosts at different rates. Our two Whisper Large V3 calls were billed at rates four times apart.

For a person with a shoebox of recordings, the honest summary is that price barely matters. A 100-hour backlog, about two years of a weekly one-hour show, costs $36 at the top of the table and $10 to $27 for the models that lead the benchmarks. Price starts to matter at product scale, when the hours run into the thousands each month. Our cost of one hour of AI post puts these numbers next to text, image and video.

A hundred hours of old recordings costs between nothing and $36 to turn into text. Illustration generated with Seedream 5 Pro via Runware.

Speaker labels and timestamps decide more than accuracy does

When every model gets 98% of the words, the deciding question becomes what comes back besides the words. Three extras matter most.

Speaker labels. Called diarization, this marks who said what. For an interview it is the difference between a transcript and a wall of text. Scribe v2 labels up to 32 speakers at no extra charge. AssemblyAI adds it for $0.02 per hour. OpenAI sells it as a separate model, GPT-4o Transcribe Diarize, and MAI-Transcribe 2 added it at launch. Open-weight Whisper has none built in. One practical warning from our test: none of the 16 models returned speaker labels through the router’s endpoint. If you need them, call the vendor’s own API.

Timestamps. Subtitles and searchable audio need to know when each word was said. Scribe v2, MAI-Transcribe 2 and Parakeet return word-level times. Whisper’s standard output times each segment. OpenAI’s documentation for GPT Transcribe does not mention timestamps at all, so a subtitle workflow on OpenAI still runs through Whisper-1.

Languages. Whisper still has the widest reach, with 99 languages on its model card. Scribe v2 covers more than 90 and MAI-Transcribe 2 covers 60. Parakeet v3 handles 25 European languages, and Canary-Qwen, the open accuracy leader, is English only with a 40-second input limit. Outside English and the major European languages, test on your own audio before trusting any leaderboard. Length limits vary too: Scribe accepts files up to 3 GB and 10 hours, which saves you from cutting a long recording into pieces.

A laptop transcribes an hour of audio in about two minutes

Whisper may have lost the accuracy crown, but it is still the model you can download. The weights carry an MIT license, and whisper.cpp runs them on a Mac, a Windows PC or a Linux box with no Python and no account. We ran the Turbo model, compressed to a 547 MB file, on the same two clips.

Four commands, no account, no upload
$ brew install whisper-cpp ffmpeg
$ curl -L -o turbo.bin https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v3-turbo-q5_0.bin
$ ffmpeg -i interview.mp3 -ar 16000 -ac 1 -c:a pcm_s16le interview.wav
$ whisper-cli -m turbo.bin -f interview.wav -otxt -osrt
Run as written on a MacBook Pro with an M3 Max and 36 GB of memory, whisper.cpp 1.9.4 from Homebrew. The 164-second conversation clip finished in 5.8 seconds, about 28 times faster than real time. The last command writes a plain text file and an .srt subtitle file next to the audio.

The local run matched NASA’s transcript of the intro word for word, and it kept every passage that the hosted copies dropped. We cannot see inside the hosts, so we cannot say why. The lesson carries over either way: “Whisper” is not one product. The same weights behave differently depending on who runs them and how they split your audio.

At 28 times real time, an hour-long interview takes a little over two minutes on that laptop. The Whisper repository lists about 6 GB of memory for Turbo in its original form, and whisper.cpp lists under 1 GB for the 466 MB “small” model, which fits on almost any computer. Owners of an NVIDIA card have a second option in Parakeet v3, which is free for commercial use under CC BY 4.0, needs only 2 GB of memory and returns word timestamps.

The strongest argument for local is not the saved cent. It is that the recording never leaves the machine. A journalist’s source, a therapy session, a client call and a family argument all make poor uploads. CSuite’s audio workspace has a Transcribe mode built on this idea: Whisper Base, Small and Large V3 Turbo run on-device, and the same screen can send a file to OpenAI, ElevenLabs or OpenRouter when you want speaker labels or a cloud model’s accuracy. We listed other jobs that work with the network off in everyday wins for offline AI.

The private option: the recording, the model and the transcript all stay on one machine. Illustration generated with Seedream 5 Pro via Runware.

Pick by the job, then check the hard parts

Podcasters and interviewers need speaker labels first. Scribe v2 includes them, with word timestamps, at $0.22 per hour. AssemblyAI comes to $0.23 with labels added. Scribe v2 ranked fourth on Artificial Analysis when we checked. If the next step is a narrated version, our AI voiceover guide covers the other direction.

Journalists, researchers and anyone with sensitive audio should start local. Whisper Turbo through whisper.cpp was flawless on our clean clip and complete on the conversation. Add speaker names by hand, which for a two-person interview takes minutes.

Students with a lecture backlog can use the same local setup for free, or a hosted open model for about a cent per hour. Lectures are one voice into one microphone, the easy case.

Developers building a product should test MAI-Transcribe 2 and GPT Transcribe against their own audio, and budget for Microsoft’s price changing in January.

Whichever you pick, do not judge it on the first clean paragraph. Every model passes that test now. Take the messiest three minutes you have, the part with two people talking and a name nobody spelled, and read the output against the audio once. Look for the sentence that is missing, because that is the error no spell checker will ever find.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app