Reviewed and written up in the last 7 days.
New
Added Sep 4, 2026
Snowflake AI Gateway is an enterprise inference routing layer released September 2026. It intelligently routes queries to the most cost-effective model that meets quality requirements. Users define acceptable accuracy/latency thresholds, and the gateway automatically chooses between Claude, GPT, Gemini, or open-source alternatives based on current pricing and performance. Uses machine learning to predict which model is optimal for each query pattern. Integrated directly into Snowflake SQL and Python notebooks.
New
Added Sep 1, 2026
Gemini 3.5 Transcribe is Google's specialized speech-to-text model released August 2026, part of the Gemini 3.5 ecosystem. Built specifically for transcription, it leverages Gemini's multimodal reasoning to improve accuracy on technical terminology, accents, and domain-specific vocabulary. Supports real-time streaming transcription and batch processing. Features confidence scoring per segment and speaker diarization (beta). API pricing based on audio duration processed. Key differentiator: uses Gemini reasoning to maintain context across long audio files for better technical accuracy.
New
Added Sep 4, 2026
Alibaba released a 7-billion parameter language model optimized for consumer laptop deployment, released August 2026. Quantized to 4-bit with custom ONNX optimization for CPU/GPU inference. Competitive response to Meta's on-device model strategy, targeting Windows/Mac laptops with 8GB+ RAM. Open weights under OpenMDW-1.1 license. Achieves reasonable performance on everyday tasks (email drafting, code generation) while running entirely offline without cloud dependency.
New
Added Sep 1, 2026
LTX-2.5 is an open-weights video and world model from Lightricks (LTX company spun out of Lightricks), released August 2026. Generates 10-second video clips from images in 6.8 seconds on Nvidia superchips. Features multi-shot support, diffusion decoder for higher quality, new conditioning modes, and autoregressive models for real-time use and robotics. Weights freely available on Hugging Face under OpenMDW-1.1 license. 33 million downloads, most-used open world model line on the market. Free for organizations under $10 million annual revenue; larger companies negotiate licenses.
New
Added Sep 4, 2026
Google Gemma is a family of open-source language models available in 2B, 7B, and 27B parameter sizes. Trained on 6 trillion tokens of high-quality data, designed for efficient local deployment. Includes standard and instruction-tuned variants. Reached 1 billion total downloads across Hugging Face, Kaggle, and GitHub by August 2026. Has spawned 100,000+ community derivatives (finetuned versions, specialized variants, quantizations). Available under Google DeepMind's permissive license for commercial and research use.
New
Added Sep 4, 2026
Cursor Origin is a new code hosting platform launched as Cursor's competitive response to GitHub. Integrates directly with Cursor IDE for seamless AI-assisted coding workflows.
New
Added Sep 4, 2026
Dulo is a new startup by Waymo pioneer Sebastian Thrun building foundation models specifically for hardware design. Uses AI to accelerate chip and hardware development cycles.
New
Added Sep 4, 2026
Grok 4.7 is SpaceXAI's reasoning-optimized model released September 12, 2026, positioned as a direct competitor to Claude Opus 5 and GPT-5.6 Sol on reasoning benchmarks. Built by Elon Musk's xAI division, Grok 4.7 features extended reasoning chains, real-time information access via X integration, and optimized inference for latency-sensitive applications. Supports API-only deployment with freemium pricing (free tier + paid enterprise).
New
Added Sep 1, 2026
Muse Glimmer is Meta's 30 billion parameter open-weight model released August 2026, designed to run autonomous AI agents directly on consumer hardware without cloud dependency. It ships under the permissive Apache 2.0 license (Meta's first fully open release since moving to proprietary Muse Spark in April), quantized to roughly 4 bits with block-level speculative decoding for fast local inference. Drops the 700M monthly user restriction that hampered prior Llama releases.
New
Added Sep 4, 2026
Claude Fable 5.1 is Anthropic's Fable-tier creative and reasoning model with integrated text watermarking. Released September 2026 as an update to Claude Fable 5, it features automatic watermarking of all text outputs plus a public detection API to verify AI-generated content. Watermarks are imperceptible to users but machine-verifiable, meeting EU AI Act transparency requirements. Priced competitively as a fast model suitable for high-volume applications.
New
Added Sep 4, 2026
Claude Mythos 5.1 is Anthropic's flagship Mythos-tier model released September 2026, representing the most advanced reasoning capabilities in the Claude lineup. Built for research, complex analysis, and multi-step problem solving, it features extended thinking chains, controlled research output, and advanced content moderation. Offers configurable reasoning depth (low/medium/high/xhigh/max) to balance accuracy and latency. Also includes optional text watermarking for transparency.
New
Added Sep 1, 2026
Laguna S 2.1 is a 118 billion parameter Mixture-of-Experts model from Poolside AI, released August 2026. Activates only 8 billion parameters per token, supports context windows up to 1 million tokens, runs under permissive OpenMDW-1.1 license. Benchmarks: 70.2% on Terminal-Bench 2.1 (beating DeepSeek-V4-Pro-Max, Nvidia Nemotron 3 Ultra), 78.5% on SWE-Bench Multilingual. Trained in under 9 weeks on 4,096 Nvidia H200 GPUs.