RADAR ARCHIVE

Week 36, 2026

17 tools added to the directory and 2 launches tracked, Aug 31 – Sep 6, 2026.

Added to the directory (17)

Tested and written up.

OpenAI's purpose-trained security model for vulnerability discovery and exploit development
New Added Sep 5, 2026
GPT-5.6-Cyber is OpenAI's security-focused variant of GPT-5.6 Sol, trained specifically for cybersecurity work including zero-day discovery, exploit chain building, and vulnerability research. Achieves 95% accuracy on advanced cybersecurity prompts compared to 1.5% for standard Sol. Deployed through OpenAI's Daybreak Red program with strict vetting and restricted access for authorized security researchers and enterprises.
Why: First frontier model purpose-trained for offensive security capabilities. Demonstrates dual-use risks of advanced AI. Important example of responsible deployment restricting high-risk capabilities.
proprietary Best for Security Research Visit
Persistent AI agents that operate your apps autonomously—$120/month
New Added Sep 6, 2026
Grok Bot by SpaceXAI is an agentic product that shifts AI assistants from answering prompts to continuously executing work inside applications employees already use. Users create persistent bots with specific jobs, grant them access to apps/websites, and delegate tasks as they would to teammates. Each bot runs in its own computer environment, keeps working when the user's laptop closes, and checks back when needing approval or finishing. Started as internal prototype for sales outbound, marketing campaigns, office operations, and bug fixes.
Why: Represents shift from AI chatbots to autonomous persistent agents. Early commercial agentic work product showing real enterprise adoption.
Paid Best for Autonomous Agents Visit
Chinese frontier model from Zhipu AI, competitive with top open-weight models, API-only
New Added Sep 4, 2026
GLM-5.3 is Zhipu AI's (z.ai) latest frontier language model, scoring 60 on Artificial Analysis Intelligence Index—seven points above GLM-5.2 and tied with Kimi K3 as the strongest open-weight model available. Sharp gains on agentic coding and cyber benchmarks including Terminal-Bench 3.0. Released as API-only (weights promised but not yet released). Pricing: $1.40/M input tokens, $4.40/M output tokens, with cached input at $0.26 and cache storage free for limited time.
Why: Demonstrates competitive strength of Chinese AI research. Positioned between open-source and proprietary models. Aggressively priced competitive with open-source APIs. Weights promised suggest eventual open-source release.
Freemium Best for Cost Efficiency Visit
Improved text-to-image generation with better detail, faster speed, and text rendering
New Added Sep 6, 2026
ChatGPT Images 2.0 is OpenAI's updated image generation model integrated into ChatGPT. Improves upon prior versions with better detail rendering, faster generation speed (up to 4x faster), and notably improved text rendering in images. Can generate images from text prompts and edit existing photos with more precision. Ranked second globally in text-to-image generation and image editing (behind previous versions of its own gpt-image-2 model). Available to ChatGPT users with Plus/Pro subscriptions.
Why: Shows iterative improvement in text-to-image quality and speed. Better text rendering is significant for design and creative use. Demonstrates practical advancement in generative image quality.
Freemium Best for Text-to-Image Generation Visit
Open-weights video and world model, generates 10-second clips in 6.8 seconds
New Added Sep 4, 2026
LTX-2.5 is an open-weights video and world model from LTX (spun out of Lightricks). Generates 10-second video clips from images in 6.8 seconds on Nvidia superchips. Features multi-shot support, diffusion decoder for higher quality, new conditioning modes, and autoregressive models for real-time use and robotics. Weights available on Hugging Face and through LTX API. Free for organizations under $10M ARR; larger companies negotiate license. Over 33 million downloads of LTX models—most-used open world model line.
Why: Open-weights video model reaching massive adoption (33M downloads). Demonstrates accessibility of frontier video capabilities. Shows viability of open-source generative models at scale.
Freemium Best for Open Video Generation Visit
Intelligent routing to optimal models for 3x cost savings on inference
New Added Sep 4, 2026
Snowflake AI Gateway is an enterprise inference routing layer released September 2026. It intelligently routes queries to the most cost-effective model that meets quality requirements. Users define acceptable accuracy/latency thresholds, and the gateway automatically chooses between Claude, GPT, Gemini, or open-source alternatives based on current pricing and performance. Uses machine learning to predict which model is optimal for each query pattern. Integrated directly into Snowflake SQL and Python notebooks.
Why: Cost optimization at scale is major enterprise concern. Automatic model selection removes guesswork and prevents overpaying for high-capability models on simple tasks. Claimed 3x savings shown in practice through intelligent model tiering. Direct Snowflake integration means no architectural changes required.
Paid Best for Cost Optimization Visit
Gemini 3.5-powered speech-to-text with contextual accuracy for technical/specialized content
New this month Added Sep 1, 2026
Gemini 3.5 Transcribe is Google's specialized speech-to-text model released August 2026, part of the Gemini 3.5 ecosystem. Built specifically for transcription, it leverages Gemini's multimodal reasoning to improve accuracy on technical terminology, accents, and domain-specific vocabulary. Supports real-time streaming transcription and batch processing. Features confidence scoring per segment and speaker diarization (beta). API pricing based on audio duration processed. Key differentiator: uses Gemini reasoning to maintain context across long audio files for better technical accuracy.
Why: Google's specialized transcription model distinct from general Gemini. Gemini-powered reasoning significantly improves accuracy on technical content vs. traditional ASR. Real-time + batch flexibility covers enterprise and consumer use cases. Emerging diarization feature and confidence scores enable quality auditing.
Freemium Best for Technical Transcription Visit
7B quantized model for offline laptop deployment, competes with Meta on-device push
New Added Sep 4, 2026
Alibaba released a 7-billion parameter language model optimized for consumer laptop deployment, released August 2026. Quantized to 4-bit with custom ONNX optimization for CPU/GPU inference. Competitive response to Meta's on-device model strategy, targeting Windows/Mac laptops with 8GB+ RAM. Open weights under OpenMDW-1.1 license. Achieves reasonable performance on everyday tasks (email drafting, code generation) while running entirely offline without cloud dependency.
Why: On-device AI becoming competitive necessity. Alibaba's direct challenge to Meta's laptop focus shows enterprise interest in consumer inference. Open weights under permissive license removes licensing friction for deployment and modification. Practical alternative for users valuing privacy and offline capability.
Free Best for On-Device Inference Visit
Open-weight video/world model, 10s clips from images in 6.8s
New this month Added Sep 1, 2026
LTX-2.5 is an open-weights video and world model from Lightricks (LTX company spun out of Lightricks), released August 2026. Generates 10-second video clips from images in 6.8 seconds on Nvidia superchips. Features multi-shot support, diffusion decoder for higher quality, new conditioning modes, and autoregressive models for real-time use and robotics. Weights freely available on Hugging Face under OpenMDW-1.1 license. 33 million downloads, most-used open world model line on the market. Free for organizations under $10 million annual revenue; larger companies negotiate licenses.
Why: LTX-2.5 dominates the open video/world model space (33M downloads). Open weights under permissive license, strong feature set (multi-shot, better quality, robotics support). For teams building video generation or world model workflows without proprietary constraints, this is the category leader.
Freemium Best for Open Video Generation Visit
Lightweight models with 1B+ downloads and thriving 100K+ derivative ecosystem
New Added Sep 4, 2026
Google Gemma is a family of open-source language models available in 2B, 7B, and 27B parameter sizes. Trained on 6 trillion tokens of high-quality data, designed for efficient local deployment. Includes standard and instruction-tuned variants. Reached 1 billion total downloads across Hugging Face, Kaggle, and GitHub by August 2026. Has spawned 100,000+ community derivatives (finetuned versions, specialized variants, quantizations). Available under Google DeepMind's permissive license for commercial and research use.
Why: 1B+ downloads demonstrates successful open-source adoption. 100K+ derivatives show strong community extending and adapting the models. Covers efficiency needs from edge (2B) to capabilities (27B). Strong proof that open models achieve massive scale in production deployments.
Free Best for Open Source Development Visit
Code hosting platform, competitive alternative to GitHub
New Added Sep 4, 2026
Cursor Origin is a new code hosting platform launched as Cursor's competitive response to GitHub. Integrates directly with Cursor IDE for seamless AI-assisted coding workflows.
Why: Cursor's move into code hosting shows vertical integration in AI coding space. Direct IDE integration removes friction in developer workflows.
Freemium Best for AI-Assisted Coding Visit
Foundation models for hardware design
New Added Sep 4, 2026
Dulo is a new startup by Waymo pioneer Sebastian Thrun building foundation models specifically for hardware design. Uses AI to accelerate chip and hardware development cycles.
Why: Novel application of foundation models to hardware design. Thrun's Waymo pedigree and focus on robotics-adjacent hardware make this significant for embodied AI.
Paid Best for Hardware Design Visit
SpaceX's reasoning model, competitive with frontier LLMs
New Added Sep 4, 2026
Grok 4.7 is SpaceXAI's reasoning-optimized model released September 12, 2026, positioned as a direct competitor to Claude Opus 5 and GPT-5.6 Sol on reasoning benchmarks. Built by Elon Musk's xAI division, Grok 4.7 features extended reasoning chains, real-time information access via X integration, and optimized inference for latency-sensitive applications. Supports API-only deployment with freemium pricing (free tier + paid enterprise).
Why: Significant new entrant in frontier LLM space backed by SpaceX capital. Reasoning capabilities competitive with established frontier models. Real-time information access via X API integration differentiates it from closed-model competitors. Early benchmarks show competitive performance on reasoning tasks.
Freemium Best for Advanced Reasoning Visit
30B on-device AI agent, runs natively on consumer hardware
New this month Added Sep 1, 2026
Muse Glimmer is Meta's 30 billion parameter open-weight model released August 2026, designed to run autonomous AI agents directly on consumer hardware without cloud dependency. It ships under the permissive Apache 2.0 license (Meta's first fully open release since moving to proprietary Muse Spark in April), quantized to roughly 4 bits with block-level speculative decoding for fast local inference. Drops the 700M monthly user restriction that hampered prior Llama releases.
Why: Open weights on Apache 2.0, runs on consumer hardware, removes user-count restrictions that hampered Llama. For teams building local-first or edge-deployed agents, this eliminates licensing friction and eliminates inference costs.
Free Best for On-Device Agents Visit
Fast reasoning model with built-in text watermarking for compliance
New Added Sep 4, 2026
Claude Fable 5.1 is Anthropic's Fable-tier creative and reasoning model with integrated text watermarking. Released September 2026 as an update to Claude Fable 5, it features automatic watermarking of all text outputs plus a public detection API to verify AI-generated content. Watermarks are imperceptible to users but machine-verifiable, meeting EU AI Act transparency requirements. Priced competitively as a fast model suitable for high-volume applications.
Why: Watermarking + detection API represents significant move toward AI-generated content provenance. EU AI Act compliance built-in from launch, critical for enterprises. Fast execution speed with transparency features addresses emerging regulatory requirements without sacrificing performance.
Paid Best for Fast Reasoning Visit
Anthropic's most advanced reasoning model for complex multi-step tasks
New Added Sep 4, 2026
Claude Mythos 5.1 is Anthropic's flagship Mythos-tier model released September 2026, representing the most advanced reasoning capabilities in the Claude lineup. Built for research, complex analysis, and multi-step problem solving, it features extended thinking chains, controlled research output, and advanced content moderation. Offers configurable reasoning depth (low/medium/high/xhigh/max) to balance accuracy and latency. Also includes optional text watermarking for transparency.
Why: Mythos-tier is Anthropic's highest capability level. Advanced reasoning and extended thinking make it ideal for research teams, scientific analysis, and complex decision support. Research mode outputs improve reproducibility and auditability for institutional use cases.
Paid Best for Advanced Research Visit
118B MoE model beats rivals 10x its size on coding benchmarks
New this month Added Sep 1, 2026
Laguna S 2.1 is a 118 billion parameter Mixture-of-Experts model from Poolside AI, released August 2026. Activates only 8 billion parameters per token, supports context windows up to 1 million tokens, runs under permissive OpenMDW-1.1 license. Benchmarks: 70.2% on Terminal-Bench 2.1 (beating DeepSeek-V4-Pro-Max, Nvidia Nemotron 3 Ultra), 78.5% on SWE-Bench Multilingual. Trained in under 9 weeks on 4,096 Nvidia H200 GPUs.
Why: Benchmark contender: claims to beat models many times its size on two critical coding benchmarks. Open weights under permissive license. Rapid training timeline (9 weeks) suggests efficient engineering. Strong SWE-Bench showing makes it worth evaluating for coding agent workloads.
Free Best for Coding Performance Visit

On the radar (2)

Reported in the daily briefing this week. Not reviewed for the directory.

VentureBeat 2026-09-03 8.7

'Welcome to the AGI era': OpenAI launches GPT-6 Astra

OpenAI officially released GPT-6 Astra, a frontier model the company claims marks the onset of artificial general intelligence (AGI) with capabilities that outperform humans on most economically valuable work. The model can navigate software autonomously witho…

TechMeme 2026-09-03 7.8

HTC rolls out its 499 Vive Eagle smart glasses with AI assistant choice

HTC launched the Vive Eagle smart glasses, priced at $499, allowing users to choose between Google's Gemini or OpenAI's ChatGPT as their AI assistant, now available in North America, Europe, and Australia. The glasses represent a major consumer push into AI-po…