Google's coding workhorse, three weeks after 3.6 Flash
New this monthAdded Aug 13, 2026
Gemini 3
Why: It is the cheapest route to a current-generation Google coding model. Introductory pricing runs at half the standard Flash rate until the end of 2026, and the published jump over 3.6 Flash is large enough to matter on exactly the agentic and web-development work Flash-tier models are usually bought for.
Z.ai's post-trained coding and agentic model on the GLM-5.2 base
New this monthAdded Aug 14, 2026
GLM-5
Why: The Terminal-Bench 3.0 score moved from 4.6% to 28.3%, DeepSWE v1.1 from 46.2% to 66.9%, and CyberGym from 77.2% to 84.5% — all on the same base as GLM-5.2. That is a real signal about post-training returns, even if most numbers are vendor-run and weights are not yet released. It is also reported as one of the fastest models in its class, at roughly 115 tokens per second.
OpenAI has opened a preview of Ultrafast, a new API service tier that runs its flagship GPT-5.6 Sol model at up to 14 times normal speed, delivering as many as 750 output tokens per second. The tier is powered by Cerebras wafer scale hardware rather than conve…
SpaceXAI, the company formerly known as xAI, released Grok 4.6 on August 12, a new flagship model tuned for long-running agents and more ambitious interactive and visual work. The company says it scores 61 on the Artificial Analysis Intelligence Index, matchin…
Google released Gemini 3.7 Flash on August 13, only three weeks after Gemini 3.6 Flash shipped, an unusually short gap that Google attributes to algorithmic gains in the model's reasoning core rather than a larger training run. Coding scores moved sharply in t…
DeepSeek released the production build of its flagship model, DeepSeek V4 Pro, ending a preview that ran nearly four months. The version, designated V4 Pro 0813, appeared on August 12 as the general availability release on OpenRouter's model page, and DeepSeek…
Nvidia released Nemotron 3.5 Lightning, a 30 billion parameter mixture of experts model that companies can download, modify and deploy for free without asking permission. Nvidia claims it delivers up to four times faster output and 30 percent faster task compl…