Skip to content
CSuite
RoundupAI ModelsNewsAug 27, 202610 min read

AI models roundup, August 2026: the promises that shipped

July's AI labs promised weights, APIs, and price cuts. August had to deliver. What actually shipped, sorted, with receipts linked.

The monthly roundup · Edition 2 · August 2026
Seven promises came due in thirty days. Five shipped.
Kimi K3 open weights · promised Jul 17Jul 27 · 1.4 TB on Hugging Face, ahead of its own deadlineKept
FLUX 3 Video open access · promised Jul 23Aug 4 · general availability, priced per secondKept
Ling-3.0-flash weights · promised Jul 23Aug 5 · MIT license; the FP8 build fits in 128 GBKept
Seedance 2.5 public API · promised Jul (beta)Aug 7 · open to developers on BytePlus and VolcanoKept
Qwen3.8-Max weights · promised Aug 3Aug 12 · the 2.4T checkpoint, under a custom licenseKept
GLM-5.3 weights · promised Aug 14Held back after a safety review; no repo as of Aug 26Waiting
Mistral's open frontier family · promised Jul (teased)Still in early access; no weights, no model cardWaiting
Every open promise we logged in the July edition, and where it stood on August 27, 2026. Each row links to a primary source in the body. Nothing is quoted from memory.

Two frontier models launched nine days apart in August, from labs an ocean apart, and arrived at the identical price: $2 per million input tokens, $6 per million output. That number is not a coincidence. It is what a month of head-to-head competition does to a sticker. And it was only half of August’s story, because the other half was labs paying bills: nearly everything July announced, teased, or promised had to actually ship.

This is the second edition of the monthly digest, built the same way as the first: every date, price, and spec below was checked against a source this week and linked in place, vendor-only benchmarks are labeled as such, and the loudest non-events get named at the end. Seventeen releases from eleven vendors made the cut for the thirty days ending August 27. Three should change decisions.

The TL;DR: three releases matter

If you buy tokens at the top tier, Grok 4.6 reset the floor. xAI’s August 12 release claims parity with the frontier on independent composite testing while keeping Grok 4.5’s $2/$6 pricing, making it the cheapest model with a credible frontier claim. If you fine-tune or self-host, Qwen3.8-Max changed what “open” includes. Alibaba’s 2.4-trillion-parameter flagship launched August 3 and then did something no Max-class Qwen had done: the actual checkpoint went up for download nine days later. If you run a workhorse tier, Gemini 3.7 Flash cut your bill in half. Google shipped it 23 days after Gemini 3.6 Flash, at half the price and with its coding score up sixteen points. Everything else is commentary on one structural fact: the promises of July, open weights above all, mostly converted into shipped artifacts in August, and prices kept falling while it happened.

The month on one card, in release order
Jul 27Kimi K3 weights 1.4 TB of the 2.8T flagship, delivered ahead of the promise.Open
Jul 31DeepSeek-V4-Flash-0731 284B/13B-active agent retrain, MIT weights on day one.Open
Jul 31Seedance 2.5 Official release; the public developer API followed on August 7.Video
Aug 3Qwen3.8-Max 2.4T MoE at $2/$6 with 1M context; weights followed mid-month.Text
Aug 4FLUX 3 Video GA: 20-second clips with native sound from $0.06 per second.Video
Aug 4Qwen-Image-3.0 July's invite-only tease went public from $0.03 per image.Image
Aug 5Ling-3.0-flash weights MIT license; BF16 at 255 GB, FP8 at 128 GB.Open
Aug 5Muse Spark 1.2 + Muse Code 1M context and repo-scale training, plus an agent framework beta.Text
Aug 5GPT Transcribe Transcription at $0.0045 per minute, 25% under whisper-1.Audio
Aug 5Think Fast 2.0 live July's speech-to-speech tease switched on at $0.08 per minute.Audio
Aug 6GPT-5.6 Luna, free default Free ChatGPT re-based; unlimited text chats the following week.Text
Aug 8Grok Imagine Image 2.0 Second on both Arena image boards; app-only, API pending.Image
Aug 12Grok 4.6 Frontier-parity claim at $2/$6, unchanged from Grok 4.5.Text
Aug 12Seed 2.1 Turbo Multimodal coding-and-agents model at $0.50/$2.50, 262K context.Text
Aug 13Gemini 3.7 Flash Half the price of its 23-day-old predecessor, big coding gains.Text
Aug 14GLM-5.3 Large agentic gains; the promised weights stayed in safety review.Text
Aug 26GLM-5.3-Flash 320B/18B multimodal MoE, MIT weights on day one, $0.15/$0.50.Open
Seventeen releases from eleven vendors in the thirty days ending August 27, 2026. “Open” marks weights you could download the day of the event. The July 27 row lands three days before our July edition published; it is included here because that edition filed it under promises.
August was a delivery month: less unveiling, more handing over what July announced. Illustration generated with Seedream 4.5 via Runware.

Text: $2 in, $6 out is the new price of the frontier

Grok 4.6 (August 12) is the cleanest expression of the trend. Same $2/$6 rates as Grok 4.5, same 500K context window, but xAI says it now matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, one point behind Claude Fable 5, with its DeepSWE coding score up from 54 to 65.9. Two honest footnotes: the parity claim compares Grok’s default effort against rivals’ maximum tiers, and the fine print reprices everything. Past 200K prompt tokens, every token in the request bills at a doubled $4/$12 rate, and cached input quietly rose 67%. Read the rate card, not the headline.

Qwen3.8-Max (August 3) picked the same sticker from the other direction: $2/$6 for a 2.4-trillion-parameter mixture-of-experts with a 1M-token context window that takes text, images, and video as input. Alibaba’s own table shows it leading on agentic computer use, with mixed coding results; independent scores did not exist at launch, so treat the table the way vendor tables deserve. The larger fact arrived on August 12, covered below: this flagship is now downloadable.

Google answered on price alone. Gemini 3.7 Flash (August 13) launched at $0.75/$3.75, half of what Gemini 3.6 Flash cost at its own launch on July 21, with DeepSWE up from 49.0% to 65.3%. The catch is a calendar: that price is introductory and doubles to $1.50/$7.50 on January 1, 2027, so any cost model built on the launch rate has a four-month shelf life. Further down the menu, ByteDance’s Seed 2.1 Turbo (August 12) offers a multimodal coding-and-agents model at $0.50/$2.50 with a 262K context, and Meta’s Muse Spark 1.2 (August 5) added a 1M-token window and repository-scale training, per the month’s dated ledger.

OpenAI spent August on distribution rather than frontier claims: GPT-5.6 Luna became the free-tier default the week of August 6, and unlimited text chats for free users followed a week later, text only, with caps intact on uploads and images. One deadline for budget owners: Anthropic’s promotional $2/$10 rate on Claude Sonnet 5 ends August 31, when it returns to $3/$15. If your pipeline standardized on that price, your invoice moves next week.

What August’s text models charge per million output tokens
First-party list prices at launch; grey bars are earlier releases kept for scale. Bar length is on a log scale.
Claude Fable 5 (June, for scale)$50
Claude Opus 5 (July, for scale)$25
Kimi K3 (July, for scale)$15
Grok 4.6$6
Qwen3.8-Max$6
Gemini 3.7 Flash$3.75
Seed 2.1 Turbo$2.50
GLM-5.3-Flash$0.50
DeepSeek-V4-Flash$0.28
Two labs an ocean apart landed on the same $6 sticker for frontier-claim models, while a 320B open model launched at $0.50. In June, the flagship rate was $50.

Image: the best new tools skipped the API

August’s most capable new image tool cannot be called from code. Grok Imagine Image 2.0 (August 8) shipped as the quality mode inside the Grok apps and site, with region-level editing, background removal, generation from up to five reference images, and noticeably better typography. It debuted second in the world on both the text-to-image and image-editing Arena boards, behind only OpenAI’s image models. The API is “planned.” Builders get to watch consumers use a model they cannot ship against yet.

Alibaba resolved one of July’s open questions by shipping. Qwen-Image-3.0 (API on August 4) went from invite-only tease to public endpoint at $0.03 per image, $0.075 for the Pro tier at 2K resolution, aimed at dense, text-heavy layouts in 12 languages. It still shipped with no benchmarks and no model card, so the honest label remains “cheap and unverified.” Housekeeping if you generate images in production: Google shut down the Imagen 4 endpoints on August 17 in favor of its Gemini-branded image model, one more reminder that image APIs retire faster than the products built on them.

Video: 30-second clips now have a public price tag

July’s video news was locked behind enterprise doors. August opened them. ByteDance’s Seedance 2.5 got its official release on July 31 and a public developer API on August 7: up to 30 seconds of synchronized audio and video in one pass, generation steered by up to 30 reference images plus video and audio references, and timestamp-level editing of an existing clip. Verified outputs top out at 1080p, so treat July’s 4K talk as roadmap. We ran the previous generation head-to-head, first takes and all, in our Seedance comparison; the 2.5 API now makes that kind of testing possible for anyone.

Black Forest Labs kept its July promise on schedule. FLUX 3 Video went generally available on August 4: clips up to 20 seconds with dialogue, sound effects, and ambient audio generated together, at $0.06 per second in draft, $0.17 at HD, and $0.29 at full HD. A 20-second draft costs $1.20, which moves AI video from “budget line” to “rounding error” for storyboarding. BFL’s own arena numbers put it ahead of Seedance 2.0 on text-to-video by 66 Elo; those are internal evaluations, and the open-weight Dev release it teased in July still has no date.

Audio: OpenAI undercut its own Whisper

The month’s audio story is one company competing with itself. GPT Transcribe (August 5) prices speech-to-text at $0.0045 per minute, roughly $0.27 per hour, 25% below the $0.006 per minute that whisper-1 has charged since 2023, with keyword hints and multi-language hints whisper never had. Transcription was already cheap; it is now cheap even by transcription’s standards, and the interesting effects are downstream, in what people index, search, and summarize once every recording defaults to having a transcript.

The other audio entry settles a July loose end. Grok Voice’s Think Fast 2.0, announced July 29 and skipped by our last edition as too fresh to judge, switched on for all Grok Voice traffic on August 5, per the same dated ledger: speech-to-speech at $0.08 per minute with 0.7-second time to first audio. Between ten-cent transcription hours and sub-second voice agents, the full voice loop keeps getting cheaper; what has not changed since July is that the best of it remains cloud-only.

Open weights: July’s promissory notes got paid, mostly

The month’s open releases are measured in drive bays: 1.4 TB for Kimi K3, roughly 2.5 TB for the Qwen3.8-Max checkpoint, 128 GB for Ling’s FP8 build. Illustration generated with Seedream 4.5 via Runware.

Our July edition ended its open-weights section warning that “a model is open when the download link works, not when the press release says so.” August is the test of that sentence, and the results were better than we expected. Moonshot paid first: Kimi K3’s weights landed on Hugging Face on July 27, ahead of the lab’s own ten-day deadline and three days before our July edition published with the promise still filed as pending. Consider this the correction: 2.8 trillion parameters, 104B active, a 1M-token context, and about 1.4 TB of files under Moonshot’s own Kimi K3 License. It is the largest open-weight release ever made.

Then the notes kept clearing. DeepSeek-V4-Flash-0731 (July 31) shipped 284B total and 13B active parameters under MIT, 166.9 GB on disk, with an agent-focused retrain that beats DeepSeek’s own larger V4-Pro preview on all nine agentic benchmarks it publishes, at $0.14/$0.28 on the hosted API. Ant Group delivered Ling-3.0-flash under MIT on August 5, two days behind its free-window schedule, in a 128 GB FP8 build. And Alibaba did the thing nobody quite believed: the full Qwen3.8-Max checkpoint went up on August 12, roughly 2.5 TB at FP8, under a custom license with revenue-gated conditions rather than Apache. Restrictions and all, a downloadable 2.4T frontier flagship did not exist a month ago.

The counterexample is instructive. Z.ai launched GLM-5.3 on August 14 with striking agentic numbers on its own benchmarks, Terminal-Bench up from 4.6 to 28.3 against its predecessor, and promised weights “in stages after a safety evaluation.” As of August 26 there is no repository. The stated reason is worth taking seriously rather than cynically: Z.ai says the model developed multi-stage exploitation planning in cybersecurity evaluations, and gating a checkpoint you cannot un-release is what a safety review is for. But the scoreboard is what it is: five of July’s seven promises cleared, and the two still waiting, GLM-5.3 and Mistral’s teased open family, both belong to the month they actually ship.

Then, one day before this edition closed, Z.ai complicated its own counterexample. GLM-5.3-Flash (August 26) is a 320B-parameter mixture-of-experts with 18B active, natively multimodal across text, image, and video input, with a 1M-token context, and its weights went up on Hugging Face under plain MIT the day of the announcement, about 306 GiB at FP8. It scores 57 on the Artificial Analysis Intelligence Index, one point above Gemini 3.7 Flash, from an API priced at $0.15/$0.50. The model had spent its first week serving traffic anonymously as “Ox Alpha,” running, per Z.ai, entirely on domestically produced Chinese chips. Day-old numbers deserve a month of scrutiny before they move your defaults, but on those numbers this is the strongest open release of the month per dollar, and it makes the gated GLM-5.3 look like the exception in Z.ai’s catalog rather than the rule.

Local-runnable: the real gift is 17 GB

Terabyte checkpoints matter for hosts and universities, not for the machine on your desk, so here is the honest sort. The one release most readers can actually run is Qwen3.8-27B (August 14): dense, Apache 2.0, about 17 GB at a 4-bit quantization, and scoring 52 on the Artificial Analysis Intelligence Index where its predecessor scored 38. That is a frontier-2025 level of intelligence in a download that fits a 24 GB GPU or a 32 GB Mac, and it replaces the older 27B as the default local Qwen the day you install it.

One step up the hardware ladder, DeepSeek-V4-Flash-0731’s 167 GB and Ling-3.0-flash’s 128 GB FP8 build both fit a 192 GB unified-memory Mac or a two-GPU workstation, with low active-parameter counts keeping generation speed pleasant; our DeepSeek local guide covers the family’s quantized builds, which land far smaller. GLM-5.3-Flash’s 306 GiB FP8 checkpoint sits just past workstation reach until smaller quantized builds land, and Kimi K3 and Qwen3.8-Max remain multi-node projects. The arithmetic for what your own machine holds is unchanged, and the RAM and VRAM guide still does it in one subtraction.

For one machine under one desk, the month’s honest shortlist is short: a 17 GB download from Alibaba and a 128 GB one from Ant. Illustration generated with Seedream 4.5 via Runware.

Skipped on purpose

What we cut, and why. GLM-5.2 Turbo (August 17): a speed refresh of the outgoing model, published three days after its own successor; if you are choosing a Z.ai model, the decision runs through GLM-5.3 or its open Flash sibling. OpenAI’s math results (August 1): ten solved open problems with machine-checkable proofs is genuine research news, and it is not a model, a product, or a price. Wan 3.0: circulated all month as imminent; no vendor post, no repository, no marketplace listing, so per this digest’s rules it does not exist yet. Mistral’s open family: still in early access with no weights and no card, same as July. We will keep writing that sentence until the download link works.

That is edition two. The September edition lands around the 30th, built the same way: primary sources, launch-day prices, and a skipped list. Two calendar notes worth carrying forward: EU enforcement of general-purpose AI obligations began applying on August 2, which is already shaping where and when some of these models launch, and Gemini 3.7 Flash’s half-price window closes December 31. Between editions, the sortable model comparison table stays current. If an August release changed which model you reach for by default, it earned its headline. For most people, exactly three did.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app