Tested and written up.
New this month
Added Sep 9, 2026
GPT-6 Astra is OpenAI's frontier model released September 2026, claimed by the company to mark the onset of artificial general intelligence (AGI). The model can navigate software autonomously without user input, working across browsers, spreadsheets, applications, and producing finished documents and workflows. Features advanced 'recurrent depth' reasoning that operates outside traditional sequential thinking patterns. Solves decades-old mathematics problems and demonstrates autonomous agent capabilities across domains.
Why: Represents a claimed inflection point in AI autonomy. Demonstrates system-level reasoning and multi-step task completion. First model to achieve 'Critical' level cybersecurity capabilities (restricted access). Raises important questions about AI development velocity and safety infrastructure.
New this month
Added Sep 8, 2026
OpenAI Ultrafast is a new API service tier that runs GPT-5.6 Sol at up to 14 times normal speed, delivering 750 output tokens per second. Powered by Cerebras wafer-scale hardware rather than conventional GPU clusters—a speed no GPU cloud has publicly matched. Designed for enterprise workloads where latency is critical: incident response, customer support, financial analysis, and e-commerce. Currently in limited preview with expanding availability.
Why: Demonstrates extreme optimization of inference speed using specialized hardware. First public 14x speedup for frontier model. Shows infrastructure differentiation beyond model capability.
New this month
Added Sep 10, 2026
ChatGPT Images 2.5 is OpenAI's latest image generation model announced September 8, 2026. Delivers sharper detail, more natural lighting and texture, better preservation of people and products in reference photos, and 50% faster generation than Images 2.0. Reliable editing across multi-turn conversations. Available in ChatGPT (all tiers), Codex, and API with two variants: Flare (high-quality, low-latency default) and Sunburst (premium control for production workflows).
Why: Represents significant iterative improvement in image quality, speed, and editing reliability. Two API variants show sophisticated tuning for different use cases. 50% speed improvement is substantial for production workflows.
New this month
Added Sep 7, 2026
HTC Vive Eagle are consumer smart glasses priced at $499, allowing users to choose between Google's Gemini or OpenAI's ChatGPT as their built-in AI assistant. Launched in North America, Europe, and Australia. Represents major entry into AI-powered wearables, competing directly with Meta's AI glasses by bringing frontier AI models to mainstream consumer hardware in always-available form factor.
Why: Consumer smart glasses with AI assistant choice represent mainstream adoption of wearable AI. Shows shift of frontier AI from phones/computers to persistent wearable interfaces.
New this month
Added Sep 8, 2026
Google Pics is an AI-powered design platform that simplifies graphic design by replacing traditional UI tools with natural language prompts, similar to generative image models. The tool automatically handles layout, typography, color coordination, and asset selection based on user descriptions. Makes professional-quality design accessible to non-designers within seconds, directly challenging Canva's dominance in accessible design and competing with Adobe's creative suite.
Why: Represents Google's entry into the accessible design space. Demonstrates shift from tool-based to prompt-based creative workflows. Early indication of how generative AI will reshape design software.
New
Added Sep 12, 2026
SWE-2 is Cognition's flagship coding model, released September 12, 2026 and built by reinforcement-learning post-training on top of Moonshot AI's 2.8-trillion-parameter open-weight Kimi K3. Cognition reports it matches Claude Fable 5.1 on FrontierCode 1.1 Main (both around 50%) at roughly 64% lower inference cost, scores 73.0% on DeepSWE 1.1 and 92.8% on Terminal-Bench 2.1, but trails Fable 5.1 and GPT-6 Astra by around 30 points on the harder Terminal-Bench 4. It ships with no open weights and no self-serve public API: it is exclusive to Devin's desktop app and CLI, with a Devin Web and Fusion rollout planned. Usage inside Devin is free through mid-October 2026; after that it requires a paid Devin plan, and separate enterprise credit-based API access is priced at $3.00 per million input...
Why: SWE-2 is a genuine data point in the trend of labs fine-tuning open-weight Chinese base models to compete with closed frontier coding models on price, and it is the model Cognition is betting Devin's entire value proposition on. The catch is access: every benchmark above is Cognition's own, unverified independently, and there is no way to use the model outside Devin's ecosystem.