AI models roundup, September 2026: the frontier got cheaper
Three new flagships in thirty days, and the best one costs less than the one it replaced. September's releases, sorted, with receipts.
On this page
On September 22, Anthropic shipped the highest-scoring model on the independent leaderboard and charged less for it than for the model it replaced. The same day, OpenAI cut the price of its two workhorse models in half. That is the shape of September 2026: the frontier moved up and the invoice moved down, at more than one lab, in the same news cycle.
This is the third edition of the monthly digest, built the way the August one was: every date, price, and spec below was checked against a source this week and linked in place, vendor-only benchmarks are labeled as such, and the loudest non-events get named at the end. Thirty-five entries from twenty-three vendors made the cut for the thirty days ending September 28, and this edition also carries two corrections to the last one. Three releases should change decisions.
The TL;DR: three releases matter
If you buy tokens at the top tier, Claude Opus 5.5 moved the ceiling and the price at once. It leads Artificial Analysis’s Intelligence Index at 58, five points clear of anything else, at $4 per million input tokens and $20 per million output, under Opus 5’s $5/$25. If you run a workhorse tier, GPT-6 Sol and Luna halved your bill. Sol is $2/$10 and Luna is $0.10/$0.50, each half the price of the GPT-5.6 model it replaces, with a 1M-token context on both. If you self-host, Xiaomi’s MiMo-V2.6-Pro is the new open-weights leader. It is a trillion-parameter mixture-of-experts under plain MIT that scores 46 on the same index, two points behind GPT-6 Sol. Everything else is commentary on one fact: three flagships landed in thirty days, and the price of intelligence kept falling while they did.
Text: the top of the leaderboard now costs $4 in, $20 out
Claude Opus 5.5 (September 22) is the month’s headline for two reasons that rarely arrive together. Anthropic’s launch post prices it at $4 and $20 per million tokens, “20% less than Opus 5,” with cache reads at $0.20, and says typical workloads cost 40% less at default settings. On Artificial Analysis’s independent index it scores 58 at maximum effort, 56 at the next tier down, and 54 at high; the next entries on the board are Claude Fable 5.1 and GPT-6 Astra at 53. Read Anthropic’s own table the way vendor tables deserve; the independent number stands on its own.
OpenAI had the busier month. GPT-6 Astra arrived on September 3 as a limited preview and reached the API the next day at $10/$50 with a 1.05M-token context; past 272K tokens of input, the model page doubles the input rate and lifts output by half. It ties Fable 5.1 at 53. Then on September 22 came GPT-6 Sol and Luna. The pricing page lists Sol at $2/$10 where GPT-5.6 Sol was $4/$20, and Luna at $0.10/$0.50 where GPT-5.6 Luna was $0.20/$1.20. Sol scores 48, level with Meta’s Muse Spark 1.3 (September 2, $1.25/$4.25).
Anthropic’s other release matters for the bill more than the board. Claude Fable 5.1 and Claude Mythos 5.1 (September 1) keep the $10/$50 rate but cut cache reads from $1 to $0.25 per million, a 75% reduction on the tokens that dominate agent loops. Mythos 5.1 is the same model with fewer safeguards, available by invite to vetted organizations. One housekeeping note from the same pricing page: Claude Sonnet 5’s $2/$10, announced at launch as introductory through August 31, is now the standard price, and the scheduled increase to $3/$15 will not occur. Our August edition told you to expect the increase; you can stop.
Further down the menu, Gemini 3.8 Flash (September 2) launched at $0.75/$3.75 through December 31, doubling on January 1, 2027, per the Gemini pricing page, the same clock 3.7 Flash runs on; its model card puts DeepSWE at 73.7%, up from 65.3%, and the independent index puts it at 41. Grok 4.7 (September 21) keeps Grok 4.6’s $2/$6 and scores 46. And DeepSeek-V4.1-Flash (September 10) is a 552B mixture-of-experts with 8B parameters active on input and 16B on output, priced at $0.30/$1.20 at peak and half that off-peak. The old V4 Flash is retired and its name now routes to the new model. If you still call gpt-3.5-turbo-instruct, OpenAI’s deprecations page shut it off today, September 28; gpt-4-turbo, o1, o3-mini and o4-mini follow on October 23. Perplexity’s Sonar chat endpoint ended yesterday, September 27, in favor of its Agent API.
Image: OpenAI took both boards, and a Qwen license went backwards
GPT Image 2.5 (September 8) ships as two models: Flare, the fast default, and Sunburst, tuned for detail and editing. Both bill by the token at the same rates, $5 per million text tokens in, $8 per million image tokens in, and $30 per million image tokens out, across six quality tiers. As of the last week of September, Sunburst and Flare hold the top two places on Arena’s text-to-image board (1424 and 1401) and its image-editing board (1526 and 1482).
Alibaba’s Qwen-Image-2.1 (weights September 14) is the month’s most capable open image release and its most awkward. The model card describes a 7B generator that takes up to ten reference images and outputs native transparency, and it ships under the Qwen Research License, which limits use to research and evaluation. Earlier Qwen-Image releases were Apache 2.0. A license can move backwards, and this one did; Ant’s LLaDA-Image (September 4), a 6.5B generator and editor under Apache 2.0, is the open alternative that did not. At the other end of the price list, Recraft V4.1 Flash (September 23) generates an image for $0.007 with a median of 1.4 seconds, by Recraft’s own clock.
First correction. Our August edition said Grok Imagine Image 2.0’s API was “planned.” xAI’s August 7 launch post says the model was available in the API as grok-imagine-image-2.0 from the start, at $0.04 an image. We were wrong.
Video: Sora’s API is gone, and editing a clip costs 3 cents a second
The month’s biggest video event was a removal. On September 24, OpenAI’s deprecations page lists the Videos API and every sora-2 and sora-2-pro snapshot as shut down, with the recommended-replacement column blank. Anything still calling that endpoint has stopped working.
What shipped instead is cheaper and more specific. Black Forest Labs’ FLUX Video Edit [fast] (September 10) takes an existing clip up to 15 seconds and a prompt, and adds, removes, restyles, or re-voices with lip sync, at $0.03 per second of output at up to 720p. On the generation side, fal’s H3 Max, a post-trained build of MiniMax’s open H3, reached general availability on September 1 with a half-price Turbo tier on September 8: $0.16 per second at 1080p, $0.08 on Turbo. And Google’s Gemini Omni 1.1 Flash, which went GA on August 27, now tops Arena’s text-to-video board at 1516 and became free inside Google Vids on September 23.
Second correction. August’s skipped list said Wan 3.0 had “no vendor post” and therefore did not exist yet. Alibaba Cloud’s own post is dated August 13 and describes 30-second single-pass generation from text, images, or documents at $0.05, $0.10, and $0.20 per second for 480p, 720p, and 1080p; general availability followed on August 24. It sits sixth on Arena’s video board. Unlike Wan 2.x, the weights are closed. We missed it, and it belongs on the record.
Audio: an hour of transcription costs a dime
Three vendors priced an hour of speech at twenty cents or less in one month. Meta’s Muse Voice Transcribe (September 1) does streaming transcription, speaker labels, and end-of-turn detection in one model across 70-plus languages, per Meta’s research post, at $0.18 an hour. Microsoft’s MAI-Transcribe-2 (September 3) entered public preview at $0.10 an hour, a limited-time rate through 2026, with a 2.0% word error rate on Artificial Analysis’s board. xAI’s Grok Voice Transcribe 2.0 (September 18) claims half the errors of 1.0 at the same $0.10 an hour for batch and $0.20 for streaming. Alibaba’s Qwen-Audio-3.1 (September 20) cut transcription prices by up to 95%.
Voices got cheaper too, with a calendar attached. Gemini 3.8 Flash TTS and Flash-Lite TTS (September 22) add prompt-based voice design across 130 languages, and the pricing page lists $0.50 per million text tokens in and $9 per million audio tokens out for Flash, $6 for Lite, through December 31, doubling on January 1. OpenAI’s GPT-Live 1 (September 10) is a full-duplex voice layer at $0.05 a minute, billed per second, with the reasoning model behind it billed separately; Google answered with Gemini 3.8 Live (September 15) at $0.018 per minute of audio out, per the Gemini changelog. In music, Suno v6 (September 9) is the first Suno family trained on licensed catalogs from Warner, BMG, and Believe; it is app-only, with no API announced. Our voice comparison predates the Gemini pair, and it predates the model that landed the morning this edition published: ElevenLabs’ Eleven v4 and v4 Turbo (September 28, per its launch post), with audio tags for laughs, whispers, and sound effects in more than 90 languages, Turbo at a vendor-measured 100 ms median inference latency, and professional voice clones back after v3 dropped them. ElevenLabs says listeners preferred it about 75% of the time in blind tests, a vendor number that gets a month of scrutiny before it moves a default.
Open weights: an MIT model sits two points behind GPT-6 Sol
August’s scoreboard closes first. Z.ai’s GLM-5.3 checkpoint went public on August 27 per its model card, ending the two-week safety hold, under a custom GLM-5.3 License: MIT in shape, plus a security review for model-as-a-service hosts above $10 billion in revenue. It scores 45. The two other promises did not clear. Mistral’s newsroom logged a €3 billion funding round in September and no weights, and FLUX 3 Action (September 23), a 7B robotics model under a non-commercial license, is the only FLUX 3 checkpoint you can download; the image and video Dev release remains undated.
The new entries moved the ceiling. Xiaomi’s MiMo-V2.6-Pro (September 22) is a 1.02-trillion-parameter mixture-of-experts with 42B active, taking text, image, video, and audio input across a 1M-token context, under plain MIT, in a 573 GB repository. On Artificial Analysis it scores 46, the highest of any open model, from a hosted API at $0.435/$0.87. Its Flash sibling is 309B with 15B active in 178 GB at $0.14/$0.28. Tencent’s Hy4 preview (August 28) put a 770B model with 49B active under Apache 2.0. DeepSeek-V4.1-Flash shipped MIT on launch day at 510 GB. Shanghai AI Lab’s Intern-S2-397B (September 13) is a 403B multimodal MoE for scientific work under Apache 2.0. And MBZUAI’s IFM released K2 Horizon, seven Apache 2.0 models from 0.9B to 375B with the full training record published alongside. Three more cleared in the same month: Nex AGI’s Nex-N2.5 family (September 8), whose 35B Mini shipped Apache 2.0 on day one while the 1.6T Max builds on DeepSeek-V4-Pro-Base; Shanghai AI Lab’s Atria-Dawn-Preview (September 11), a 753B MIT agent model on the GLM-5.2 base; and Yandex’s AliceAI-Foundation-80B-A3B-Base (weights September 12, announced September 21), an Apache 2.0 base model with 3B active parameters.
Local-runnable: a 27B in 6 GB, with one catch
Trillion-parameter checkpoints are for hosts, so here is the honest sort for one machine. The release most readers can actually run is Prism ML’s Ternary Bonsai 2 27B (September 17): Qwen3.8-27B compressed to 1.76 bits per weight, a 5.9 GB file under Apache 2.0 with a 262K context, retaining “98.2% of aggregate benchmark performance” by the vendor’s measurement, at 143 tokens per second on an RTX 5090 and 46.8 on an M5 Max. August’s local pick was the same model at 17 GB in 4-bit. This is a third of that. The catch is on the model card: stock llama.cpp will not run these files. You need Prism’s fork or the MLX build, so the apps you already use will not load it yet.
For the smallest machines, OpenBMB’s MiniCPM5-2B (September 7, Apache 2.0) passed 800,000 downloads in three weeks, and Tencent’s AuK (September 9, MIT) puts a 1.5B speech generation and editing model on Apple Silicon and CPUs. One more takes the same bet on a custom runtime: Edge0-35B-A3B-preview (September 8, Apache 2.0) streams a 35B mixture-of-experts from SSD and claims “under 3 GiB of active memory.” Two more fit without a new runtime. Xiaomi’s MiMo-V2.6-Distill-Qwen-9B (September 21) is a 9B vision-capable distillation under MIT at 19 GB in BF16, with a vendor-reported SWE-Bench Pro of 44.6 against 32.0 for its base. One rung up, GLM-5.3-Flash, the 320B MIT release from late August, now has 1-bit builds at 93 GB, which puts a 320B model on a 128 GB Mac, and Ant’s Ling-3.0-flash-VL has an official int4 at 76 GB. Qwen-Image-2.1 quantizes to 3.9 GB and runs on any 12 GB card, research license and all. The arithmetic for what your machine holds has not changed, and the RAM and VRAM guide still does it in one subtraction.
Skipped on purpose
What we cut, and why. Kimi K2.8 Preview (September 11): closed, absent from Moonshot’s own pricing page, no model card. GLM-5.3-Prime and Qwen3.8-Max-Prime (September 23): faster serving of weights already covered, at higher prices. Fireworks Ember-1 (September 23): a two-week research preview of a retrained Kimi K3. Gemini 4: Google said on September 24 that it is in post-training. GPT-6 Cyber: a Fortune report from anonymous sources, pointing at DevDay on September 29. Claude Sonnet 5.5 and Haiku 5.5: the Opus 5.5 post says both “will follow in the coming weeks,” with no id, price, or card yet, so they belong to the month they ship. Grok 4.8, Kling 4.0: leaks and rumors, no vendor post. Siri AI (September 14): Apple’s beta runs on new Apple Foundation Models “custom-built in collaboration with Google and its Gemini models,” English first, not in the EU or China; with no API or card, it is a product story. Jev and CLM-8B: TypeSafe’s typed-decision model (September 15, $0.042 per million input tokens, waitlist) and Stanford and NVIDIA’s open counterpart (September 21, Apache 2.0) return a scored choice instead of text; a new category this ledger cannot rank yet. World models: World Labs’ Atlas, PixVerse’s R2, and Odyssey-3 all announced in September with no published price and no weights. Mistral’s open family: the same sentence as July and August. We will keep writing it until the download link works.
That is edition three. The October edition lands around the 29th, built the same way. Five dates to carry forward: September 30, when Amazon Bedrock retires Nova Canvas and Nova Reel; October 23, when OpenAI retires gpt-4-turbo, o1, o3-mini and o4-mini; December 31, when Gemini 3.7 Flash, 3.8 Flash and both 3.8 TTS models double in price; February 26, 2027, when whisper-1 and the gpt-4o transcription models shut down; and DevDay tomorrow, which will open the next window. Between editions, the sortable model comparison table stays current. If a September release changed which model you reach for by default, it earned its headline. For most people, exactly three did.


