AI models roundup, August 2026: the promises that shipped
July's AI labs promised weights, APIs, and price cuts. August had to deliver. What actually shipped, sorted, with receipts linked.
Two frontier models launched nine days apart in August, from labs an ocean apart, and arrived at the identical price: $2 per million input tokens, $6 per million output. That number is not a coincidence. It is what a month of head-to-head competition does to a sticker. And it was only half of August’s story, because the other half was labs paying bills: nearly everything July announced, teased, or promised had to actually ship.
This is the second edition of the monthly digest, built the same way as the first: every date, price, and spec below was checked against a source this week and linked in place, vendor-only benchmarks are labeled as such, and the loudest non-events get named at the end. Seventeen releases from eleven vendors made the cut for the thirty days ending August 27. Three should change decisions.
The TL;DR: three releases matter
If you buy tokens at the top tier, Grok 4.6 reset the floor. xAI’s August 12 release claims parity with the frontier on independent composite testing while keeping Grok 4.5’s $2/$6 pricing, making it the cheapest model with a credible frontier claim. If you fine-tune or self-host, Qwen3.8-Max changed what “open” includes. Alibaba’s 2.4-trillion-parameter flagship launched August 3 and then did something no Max-class Qwen had done: the actual checkpoint went up for download nine days later. If you run a workhorse tier, Gemini 3.7 Flash cut your bill in half. Google shipped it 23 days after Gemini 3.6 Flash, at half the price and with its coding score up sixteen points. Everything else is commentary on one structural fact: the promises of July, open weights above all, mostly converted into shipped artifacts in August, and prices kept falling while it happened.
Text: $2 in, $6 out is the new price of the frontier
Grok 4.6 (August 12) is the cleanest expression of the trend. Same $2/$6 rates as Grok 4.5, same 500K context window, but xAI says it now matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, one point behind Claude Fable 5, with its DeepSWE coding score up from 54 to 65.9. Two honest footnotes: the parity claim compares Grok’s default effort against rivals’ maximum tiers, and the fine print reprices everything. Past 200K prompt tokens, every token in the request bills at a doubled $4/$12 rate, and cached input quietly rose 67%. Read the rate card, not the headline.
Qwen3.8-Max (August 3) picked the same sticker from the other direction: $2/$6 for a 2.4-trillion-parameter mixture-of-experts with a 1M-token context window that takes text, images, and video as input. Alibaba’s own table shows it leading on agentic computer use, with mixed coding results; independent scores did not exist at launch, so treat the table the way vendor tables deserve. The larger fact arrived on August 12, covered below: this flagship is now downloadable.
Google answered on price alone. Gemini 3.7 Flash (August 13) launched at $0.75/$3.75, half of what Gemini 3.6 Flash cost at its own launch on July 21, with DeepSWE up from 49.0% to 65.3%. The catch is a calendar: that price is introductory and doubles to $1.50/$7.50 on January 1, 2027, so any cost model built on the launch rate has a four-month shelf life. Further down the menu, ByteDance’s Seed 2.1 Turbo (August 12) offers a multimodal coding-and-agents model at $0.50/$2.50 with a 262K context, and Meta’s Muse Spark 1.2 (August 5) added a 1M-token window and repository-scale training, per the month’s dated ledger.
OpenAI spent August on distribution rather than frontier claims: GPT-5.6 Luna became the free-tier default the week of August 6, and unlimited text chats for free users followed a week later, text only, with caps intact on uploads and images. One deadline for budget owners: Anthropic’s promotional $2/$10 rate on Claude Sonnet 5 ends August 31, when it returns to $3/$15. If your pipeline standardized on that price, your invoice moves next week.
Image: the best new tools skipped the API
August’s most capable new image tool cannot be called from code. Grok Imagine Image 2.0 (August 8) shipped as the quality mode inside the Grok apps and site, with region-level editing, background removal, generation from up to five reference images, and noticeably better typography. It debuted second in the world on both the text-to-image and image-editing Arena boards, behind only OpenAI’s image models. The API is “planned.” Builders get to watch consumers use a model they cannot ship against yet.
Alibaba resolved one of July’s open questions by shipping. Qwen-Image-3.0 (API on August 4) went from invite-only tease to public endpoint at $0.03 per image, $0.075 for the Pro tier at 2K resolution, aimed at dense, text-heavy layouts in 12 languages. It still shipped with no benchmarks and no model card, so the honest label remains “cheap and unverified.” Housekeeping if you generate images in production: Google shut down the Imagen 4 endpoints on August 17 in favor of its Gemini-branded image model, one more reminder that image APIs retire faster than the products built on them.
Video: 30-second clips now have a public price tag
July’s video news was locked behind enterprise doors. August opened them. ByteDance’s Seedance 2.5 got its official release on July 31 and a public developer API on August 7: up to 30 seconds of synchronized audio and video in one pass, generation steered by up to 30 reference images plus video and audio references, and timestamp-level editing of an existing clip. Verified outputs top out at 1080p, so treat July’s 4K talk as roadmap. We ran the previous generation head-to-head, first takes and all, in our Seedance comparison; the 2.5 API now makes that kind of testing possible for anyone.
Black Forest Labs kept its July promise on schedule. FLUX 3 Video went generally available on August 4: clips up to 20 seconds with dialogue, sound effects, and ambient audio generated together, at $0.06 per second in draft, $0.17 at HD, and $0.29 at full HD. A 20-second draft costs $1.20, which moves AI video from “budget line” to “rounding error” for storyboarding. BFL’s own arena numbers put it ahead of Seedance 2.0 on text-to-video by 66 Elo; those are internal evaluations, and the open-weight Dev release it teased in July still has no date.
Audio: OpenAI undercut its own Whisper
The month’s audio story is one company competing with itself. GPT Transcribe (August 5) prices speech-to-text at $0.0045 per minute, roughly $0.27 per hour, 25% below the $0.006 per minute that whisper-1 has charged since 2023, with keyword hints and multi-language hints whisper never had. Transcription was already cheap; it is now cheap even by transcription’s standards, and the interesting effects are downstream, in what people index, search, and summarize once every recording defaults to having a transcript.
The other audio entry settles a July loose end. Grok Voice’s Think Fast 2.0, announced July 29 and skipped by our last edition as too fresh to judge, switched on for all Grok Voice traffic on August 5, per the same dated ledger: speech-to-speech at $0.08 per minute with 0.7-second time to first audio. Between ten-cent transcription hours and sub-second voice agents, the full voice loop keeps getting cheaper; what has not changed since July is that the best of it remains cloud-only.
Open weights: July’s promissory notes got paid, mostly
Our July edition ended its open-weights section warning that “a model is open when the download link works, not when the press release says so.” August is the test of that sentence, and the results were better than we expected. Moonshot paid first: Kimi K3’s weights landed on Hugging Face on July 27, ahead of the lab’s own ten-day deadline and three days before our July edition published with the promise still filed as pending. Consider this the correction: 2.8 trillion parameters, 104B active, a 1M-token context, and about 1.4 TB of files under Moonshot’s own Kimi K3 License. It is the largest open-weight release ever made.
Then the notes kept clearing. DeepSeek-V4-Flash-0731 (July 31) shipped 284B total and 13B active parameters under MIT, 166.9 GB on disk, with an agent-focused retrain that beats DeepSeek’s own larger V4-Pro preview on all nine agentic benchmarks it publishes, at $0.14/$0.28 on the hosted API. Ant Group delivered Ling-3.0-flash under MIT on August 5, two days behind its free-window schedule, in a 128 GB FP8 build. And Alibaba did the thing nobody quite believed: the full Qwen3.8-Max checkpoint went up on August 12, roughly 2.5 TB at FP8, under a custom license with revenue-gated conditions rather than Apache. Restrictions and all, a downloadable 2.4T frontier flagship did not exist a month ago.
The counterexample is instructive. Z.ai launched GLM-5.3 on August 14 with striking agentic numbers on its own benchmarks, Terminal-Bench up from 4.6 to 28.3 against its predecessor, and promised weights “in stages after a safety evaluation.” As of August 26 there is no repository. The stated reason is worth taking seriously rather than cynically: Z.ai says the model developed multi-stage exploitation planning in cybersecurity evaluations, and gating a checkpoint you cannot un-release is what a safety review is for. But the scoreboard is what it is: five of July’s seven promises cleared, and the two still waiting, GLM-5.3 and Mistral’s teased open family, both belong to the month they actually ship.
Then, one day before this edition closed, Z.ai complicated its own counterexample. GLM-5.3-Flash (August 26) is a 320B-parameter mixture-of-experts with 18B active, natively multimodal across text, image, and video input, with a 1M-token context, and its weights went up on Hugging Face under plain MIT the day of the announcement, about 306 GiB at FP8. It scores 57 on the Artificial Analysis Intelligence Index, one point above Gemini 3.7 Flash, from an API priced at $0.15/$0.50. The model had spent its first week serving traffic anonymously as “Ox Alpha,” running, per Z.ai, entirely on domestically produced Chinese chips. Day-old numbers deserve a month of scrutiny before they move your defaults, but on those numbers this is the strongest open release of the month per dollar, and it makes the gated GLM-5.3 look like the exception in Z.ai’s catalog rather than the rule.
Local-runnable: the real gift is 17 GB
Terabyte checkpoints matter for hosts and universities, not for the machine on your desk, so here is the honest sort. The one release most readers can actually run is Qwen3.8-27B (August 14): dense, Apache 2.0, about 17 GB at a 4-bit quantization, and scoring 52 on the Artificial Analysis Intelligence Index where its predecessor scored 38. That is a frontier-2025 level of intelligence in a download that fits a 24 GB GPU or a 32 GB Mac, and it replaces the older 27B as the default local Qwen the day you install it.
One step up the hardware ladder, DeepSeek-V4-Flash-0731’s 167 GB and Ling-3.0-flash’s 128 GB FP8 build both fit a 192 GB unified-memory Mac or a two-GPU workstation, with low active-parameter counts keeping generation speed pleasant; our DeepSeek local guide covers the family’s quantized builds, which land far smaller. GLM-5.3-Flash’s 306 GiB FP8 checkpoint sits just past workstation reach until smaller quantized builds land, and Kimi K3 and Qwen3.8-Max remain multi-node projects. The arithmetic for what your own machine holds is unchanged, and the RAM and VRAM guide still does it in one subtraction.
Skipped on purpose
What we cut, and why. GLM-5.2 Turbo (August 17): a speed refresh of the outgoing model, published three days after its own successor; if you are choosing a Z.ai model, the decision runs through GLM-5.3 or its open Flash sibling. OpenAI’s math results (August 1): ten solved open problems with machine-checkable proofs is genuine research news, and it is not a model, a product, or a price. Wan 3.0: circulated all month as imminent; no vendor post, no repository, no marketplace listing, so per this digest’s rules it does not exist yet. Mistral’s open family: still in early access with no weights and no card, same as July. We will keep writing that sentence until the download link works.
That is edition two. The September edition lands around the 30th, built the same way: primary sources, launch-day prices, and a skipped list. Two calendar notes worth carrying forward: EU enforcement of general-purpose AI obligations began applying on August 2, which is already shaping where and when some of these models launch, and Gemini 3.7 Flash’s half-price window closes December 31. Between editions, the sortable model comparison table stays current. If an August release changed which model you reach for by default, it earned its headline. For most people, exactly three did.


