Skip to content
CSuite
RoundupAI ModelsNewsJuly 30, 20269 min read

AI models roundup, July 2026: the three that matter

Eighteen AI releases in thirty days. Three should change your decisions. Edition one of the monthly digest sorts them, receipts linked.

The monthly roundup · Edition 1 · July 2026
Eighteen releases landed in thirty days. Three should change your decisions.
July 24
Claude Opus 5
Near-flagship intelligence at half the flagship's price.
July 15
Inkling
975B open weights, downloadable on day one.
July 17
Kimi K3
A 2.8T model landing fourth among all frontiers.
Each dot is a model release we tracked, July 1–30, 2026; violet marks the three verdict-changers. Every date, price, and spec in this digest links to a primary source in the body. Nothing is quoted from memory.

On July 21, five models arrived in a single day: three Geminis from Google, an open-weight coding model from poolside, and an invite-only image model from Alibaba. That was a Tuesday. The rest of the month kept roughly that pace. By our count, twelve vendors shipped eighteen notable models or major updates between July 1 and July 30, 2026, across text, image, video, and audio.

Nobody needs eighteen verdicts, so this digest sorts instead of piling on. It is the first edition of a monthly series, built the same way each time: every date, price, and spec below was checked against a primary source this week and linked in place, and when a claim comes only from a vendor’s own benchmark, we say so. Three July releases genuinely change decisions. The rest get a paragraph. A few loudly promoted things get skipped, on purpose, at the end.

The TL;DR: three releases matter

If you buy AI by the token, Claude Opus 5 repriced your default. Anthropic’s July 24 release delivers close to the intelligence of Claude Fable 5, its frontier flagship, at half the price: $5 per million input tokens and $25 per million output. If you fine-tune or self-host, Inkling is the new base to evaluate. Thinking Machines Lab’s first-ever model is a 975B-parameter mixture-of-experts with weights downloadable the day it was announced, the largest US open-weight release to date. If you track where the frontier is, Kimi K3 moved it. Moonshot’s 2.8-trillion-parameter model landed fourth among all frontier models on independent testing, above every Western open model, from a lab charging less than a third of the leader’s rates. Everything else in July, including some very loud announcements, is commentary on those three facts: the frontier got cheaper, open got bigger, and the gap between them got thinner.

The month on one card, in release order
Jul 7Muse Image Meta's first image model, shipped inside Meta AI and ads tooling.Image
Jul 8Grok 4.5 SpaceXAI's coding-and-agents flagship at $2/$6 per million tokens.Text
Jul 9Muse Spark 1.1 Agentic update, plus Meta's first paid model API.Text
Jul 9GPT-5.6 rollout Sol, Terra, and Luna reached the broad ChatGPT public.Text
~JulSeedance 2.5 Native 30-second, 4K clips; phased enterprise rollout all month.Video
Jul 15Inkling 975B-total MoE, weights on Hugging Face at launch.Open
Jul 17Kimi K3 2.8T parameters, 1M context, fourth on the Intelligence Index.Text
Jul 20Qwen-Audio-3.0-TTS Took the #1 TTS arena spot at a third of incumbent prices.Audio
Jul 21Gemini 3.6 Flash Workhorse refresh: cheaper output, 17% fewer tokens per task.Text
Jul 21Gemini 3.5 Flash-Lite High-throughput tier at $0.30/$2.50, 350 tokens/second.Text
Jul 21Gemini 3.5 Flash Cyber Security-specialized variant for finding and fixing vulnerabilities.Text
Jul 21Laguna S 2.1 118B/8B-active coding MoE; the month's best laptop-adjacent release.Open
Jul 21Qwen-Image-3.0 Invite-only API with big long-prompt claims, no public benchmarks.Image
Jul 23FLUX 3 One model for image, video, audio, and robot actions; gated early access.Multi
Jul 23Grok STT 1.0 Transcription at $0.10 per audio hour, diarization included.Audio
Jul 24Claude Opus 5 Near-Fable 5 intelligence at $5/$25, the frontier repriced.Text
Jul 27Qwen3.7 Flash 1M-context vision-language endpoint at $0.03/$0.13.Text
Jul 27Ling-3.0-flash 124B/5.1B-active MoE; free API now, weights promised in August.Text
Eighteen releases from twelve vendors, July 2026. “Open” marks weights you could download the day of the announcement. Seedance 2.5’s date is approximate because its rollout was phased across the month.
The AI news cycle now produces a stack like this every month. The useful service is no longer reporting the releases; it is sorting them. Photo by Utsav Srestha on Unsplash.

Text: the price of frontier intelligence got cut in half

The month’s biggest release was also its least surprising shape: a better model for less money. Claude Opus 5 (July 24) keeps the $5/$25 pricing of the Opus 4.8 it replaces while claiming, on Anthropic’s own numbers, results within 0.5% of Fable 5 on CursorBench 3.2 at half the cost. Vendor benchmarks flatter, as always. But the structural fact needs no benchmark: the near-frontier tier now costs half of what the frontier tier did in June, and it became the default on Anthropic’s consumer plans the same day. If your team standardized on a model earlier this year, this is the release that justifies re-running your own evals.

The rest of the top tier spent July competing on price too. Grok 4.5 (July 8), the first flagship since SpaceX absorbed xAI, is pitched by Elon Musk as “an Opus-class model, but faster, more token-efficient and lower cost,” at $2/$6 per million tokens. Unusually, it was trained in part on real developer sessions from Cursor, and it is not yet available in the EU, which matters if your compliance team reads model cards. Kimi K3 (July 17) is the audacious one: 2.8 trillion total parameters, a 1M-token context window, and $3/$15 pricing, with independent testing placing it fourth of 189 models on the Artificial Analysis Intelligence Index. Moonshot promised open weights within ten days of launch; more on that promise below.

Google’s July was about the workhorse tier. Gemini 3.6 Flash (July 21) cut output pricing from $9 to $7.50 per million tokens while using 17% fewer output tokens per task, and the two stack: output spend falls by roughly 30%, not the 17% the sticker suggests. Its coding score on DeepSWE jumped from 37% to 49%. It shipped alongside a cheap high-throughput tier (3.5 Flash-Lite, $0.30/$2.50) and a security-specialized variant (3.5 Flash Cyber), plus one carefully placed sentence confirming Gemini 4 has started pre-training.

Two more text entries round out the month. Meta’s Muse Spark 1.1 (July 9) improves agentic and computer-use work, but the real news is the Meta Model API: Meta selling tokens directly, at $1.25/$4.25 per million, for the first time. And Alibaba closed the month with Qwen3.7 Flash (July 27), a vision-language endpoint with a 1M-token context window at $0.03/$0.13, the cheapest 1M-context multimodal API on the market. OpenAI’s contribution to July was distribution rather than novelty: the GPT-5.6 family began its broad public rollout on July 9; our launch-week breakdown covers who should use Sol, Terra, and Luna.

What July’s text models charge per million output tokens
First-party list prices at launch. Bar length is on a log scale; on a linear one, the bottom two bars would be invisible.
Claude Fable 5 (June, for scale)$50
Claude Opus 5$25
Kimi K3$15
Gemini 3.6 Flash$7.50
Grok 4.5$6
Muse Spark 1.1$4.25
Laguna S 2.1$0.20
Qwen3.7 Flash$0.13
The spread from Opus 5 to Qwen3.7 Flash is 192x. The models are not interchangeable, but in July every tier of the market got cheaper, and the top tier got cheaper fastest.

Image: distribution shipped, capability teased

July’s image story was about where models live, not how good they are. Muse Image (July 7), the first image model from Meta Superintelligence Labs, generates and edits photos inside Meta AI, with rollout planned across Facebook, Instagram, WhatsApp, and Meta’s ad tooling. Free daily creation, subscriptions for volume. As a piece of image research it did not move any leaderboard; as a distribution event it puts a current-generation image model in front of billions of people who will never visit an AI website.

The capability tease came from Alibaba: Qwen-Image-3.0 (July 21) arrived as an invite-only API claiming it can follow 4,500-token prompts and render legible 10-pixel text, per release-week coverage. No public benchmarks accompany the claims, so the honest verdict is “promising, unverifiable.” Nothing released in July displaced the incumbents at the top of the image leaderboards; if you picked an image stack in June, July gives you no reason to change it.

Video: the 30-second ceiling broke, for a price

For two years, AI video has meant clips of five to fifteen seconds. ByteDance’s Seedance 2.5, announced at its Volcano Engine conference and rolled out through July, generates native 30-second videos at up to 4K, with pricing reported above $28 per 30-second 4K clip. That is a professional tool at a professional price, aimed at the film and advertising pipelines where Seedance 2.0 already dominates China’s micro-drama industry. It spent the month in staged enterprise beta, so treat it as real but not yet general. We tested this family’s previous generation ourselves, first takes and all, in our Seedream and Seedance post.

The more architecturally interesting video news was FLUX 3 (announced July 23): Black Forest Labs’ first frontier model trained jointly on images, video, audio, and robot actions in one architecture. Video generation, up to 20 seconds with native sound, opened in gated early access; the image variant lands “in coming weeks”; an open-weight Dev release is promised later in 2026. BFL says early evaluations lead frontier video rivals, but those are vendor-run numbers with full results not yet published, so file them as marketing until the model is public. The direction is the takeaway: the image labs are converging on single models that see, hear, and act.

Audio: voice joined the price war

The clearest price-war casualty in July was text-to-speech. Qwen-Audio-3.0-TTS (July 20) took the #1 spot on the independent Artificial Analysis speech arena at roughly 1,236 Elo, a statistical tie with the 1,234 of runner-up Simba 3.2, while charging $27.59 per million characters, roughly a third of incumbent rates for the tiers it outranks, per launch-week coverage. It speaks 16 languages, clones voices, and follows natural-language style instructions. Two caveats before you migrate an audiobook pipeline: it is hosted-only on Alibaba Cloud with no downloadable weights, and the quality-focused Plus tier generates slowly, so it suits batch narration rather than live agents.

Transcription got the same treatment from the other direction. Grok STT 1.0 (July 23) prices speech-to-text at $0.10 per audio hour for batch work, with word-level timestamps and speaker diarization included. At ten cents an hour, transcribing everything, every meeting, every interview, every voicemail, stops being a budgeting decision and becomes a default. Between cheap ears and cheap voices, July made the full voice loop nearly free; what it did not change is that the best of both remain cloud-only.

Open weights: bigger than ever, downloadable later

“Open weights” increasingly means hardware like this. July’s open releases average hundreds of billions of parameters; only one fits anything you would carry in a bag. Photo by Taylor Vick on Unsplash.

The month’s most consequential open release came from a lab shipping its first model ever. Inkling (July 15), from Mira Murati’s Thinking Machines Lab, is a mixture-of-experts model with 975B total parameters and 41B active, trained on 45 trillion tokens of text, images, audio, and video, with a context window up to 1M tokens. The weights went up on Hugging Face the day of the announcement. The lab is refreshingly plain that Inkling is “not the strongest overall model available today”; it is built as a broad base for fine-tuning, and a preview of Inkling-Small (276B total, 12B active) shipped alongside it. It is the largest US open-weight release to date, and the first credible American answer to a question China’s labs have owned for a year.

The sharpest open release was smaller and pointier. Laguna S 2.1 (July 21) from poolside packs 118B total parameters, 8B active, into a coding model that scores 70.2% on Terminal-Bench 2.1 and tops the published SWE-Bench Multilingual table at 78.5%, matching models several times its size. Weights shipped on Hugging Face day one under the OpenMDW-1.1 license, with a hosted endpoint at $0.10/$0.20 for the lazy. We walked through running its predecessor on a laptop in our Laguna local-run guide; this release strengthens that path rather than replacing it.

And then there is the pattern to watch: “open” announced, download pending. Kimi K3 launched July 17 with weights promised within ten days. Ant Group’s Ling-3.0-flash (July 27), an efficiency play with 124B total parameters and just 5.1B active, is free to use through August 3, with weights promised after the promotional window. FLUX 3’s Dev release is “later in 2026.” None of these promises is necessarily hollow, and Moonshot and Ant both have shipping histories. But a model is open when the download link works, not when the press release says so. Until then, plan around these as closed models with a promissory note attached.

Local-runnable: mostly big iron, one real gift

Honest verdict for the run-it-yourself crowd: July was a big-iron month. Kimi K3’s weights are estimated at roughly 1.4 TB even in compressed MXFP4 form, with Moonshot recommending a 64-accelerator node to serve it. Inkling’s 975B parameters put it firmly in multi-GPU territory, and even Inkling-Small, at 276B total, is a workstation project. These releases matter enormously for the ecosystem, because hosts, universities, and enterprises can now run near-frontier models on their own terms. They will not run on your laptop.

The exception is Laguna S 2.1. At 118B total parameters with a mixture-of-experts design, a 4-bit quantization lands in the 60–70 GB range: poolside says it runs on a single DGX Spark workstation, and a 128 GB unified-memory Mac clears the bar comfortably. If Ling-3.0-flash’s weights arrive in August as promised, its 124B/5.1B-active shape should behave similarly, with the low active count keeping token speed pleasant on unified memory. For the subtraction that tells you what your own machine can hold, our RAM and VRAM guide still applies unchanged.

The laptop crowd got exactly one July release worth planning around, and it came from poolside, not from a trillion-parameter lab. Photo by Ales Nesetril on Unsplash.

Skipped on purpose

A digest earns trust by what it leaves out, so here is what we cut, each with the reason. Qwen3.8-Max-Preview (July 19): a preview endpoint whose “second only to the leader” claim has no public benchmarks behind it; we will cover the real release. Muse Video: widely written up alongside Muse Image as if it shipped; Meta’s own announcement says it is still in development. It is not a release. Gemini 4: one sentence about pre-training progress is a teaser, not a model. Mistral’s new open-weight family: entered early access in July with no weights and no model card; it belongs to the month it actually ships. Grok Voice Think Fast 2.0 (July 29): landed the day before we published, too late to evaluate honestly. It gets its hearing in August.

That is edition one. The August edition lands around the 30th, built the same way: primary sources, launch-day prices, and a skipped list. Between editions, the sortable model comparison table stays current with the five numbers that matter per model. If a July release changed which model you reach for by default, it earned its headline. For most people, exactly three did.

More reading
Launch offer · 50% off

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

$98$49
Pricing

Secure checkout via Stripe. Already have a license? Download the app