AI Model Release Tracker
Curated foundation model and API releases across modalities. Updated as new drops are added to the directory.
September 2026 (10)
Snowflake AI Gateway
Sep 4, 2026Intelligent routing to optimal models for 3x cost savings on inference
Gemini 3.5 Transcribe
Sep 1, 2026Gemini 3.5-powered speech-to-text with contextual accuracy for technical/specialized content
Alibaba Laptop-Ready Model
Sep 4, 20267B quantized model for offline laptop deployment, competes with Meta on-device push
LTX Video 2.5
Sep 1, 2026Open-weight video/world model, 10s clips from images in 6.8s
Google Gemma
Sep 4, 2026Lightweight models with 1B+ downloads and thriving 100K+ derivative ecosystem
Grok 4.7
Sep 4, 2026SpaceX's reasoning model, competitive with frontier LLMs
Meta Muse Glimmer
Sep 1, 202630B on-device AI agent, runs natively on consumer hardware
Claude Fable 5.1
Sep 4, 2026Fast reasoning model with built-in text watermarking for compliance
Claude Mythos 5.1
Sep 4, 2026Anthropic's most advanced reasoning model for complex multi-step tasks
Laguna S 2.1
Sep 1, 2026118B MoE model beats rivals 10x its size on coding benchmarks
August 2026 (13)
Seedance 2.5
Aug 4, 202630-second 4K video with native audio and up to 50 reference inputs
Claude Opus 5
Aug 4, 2026Anthropic's frontier model, currently first on the Artificial Analysis Intelligence Index
OpenArt
Aug 8, 2026Node workflows without running your own GPU
Weavy
Aug 8, 2026One canvas, many models, wired together
Invoke
Aug 8, 2026Open-source node canvas built around the edit, not the prompt
Ox Alpha
Aug 20, 2026Anonymous 1M-context reasoning model available free through OpenRouter
Gemini 3.7 Flash
Aug 13, 2026Google's coding workhorse, three weeks after 3.6 Flash
GLM 5.3
Aug 14, 2026Z.ai's post-trained coding and agentic model on the GLM-5.2 base
Muse Spark 1.1
Aug 4, 2026Meta's closed-weight agentic model, and its first paid model API
NVIDIA Nemotron 3 Ultra
Aug 4, 2026NVIDIA's 550B open-weights reasoning model, built for inference speed
FLUX 3
Aug 4, 2026Multimodal model generating image, video and audio from one set of weights
GLM-5.2
Aug 4, 2026753B open-weight MoE coding model with a 1M-token context, MIT licensed
Inkling
Aug 4, 2026Thinking Machines' 975B Apache-2.0 model that takes text, images and audio natively
July 2026 (18)
Claude Fable 5
Jul 7, 2026Anthropic's Mythos-class creative model
GPT-5.6 Sol
Jul 9, 2026OpenAI's top-tier model for complex professional work
Kimi K3
Jul 16, 2026Moonshot AI's 2.8-trillion-parameter open-weight flagship
Grok 4.5
Jul 9, 2026xAI's flagship coding model, trained in partnership with Cursor
MiniMax H3
Jul 31, 2026Omni-modal video with native stereo audio, at 2K
OpenCode
Jul 7, 2026Open-source, model-agnostic terminal coding agent
Kimi K2.7-Code
Jul 7, 2026Moonshot's specialized coding model
Smart Topology
Jul 21, 2026Clean, controllable game-ready topology in ~10 seconds
Meshy 3D Agent
Jul 21, 2026Conversational AI agent for end-to-end 3D creation
Luma Ray 3.2
Jul 15, 2026Controllable cinematic video model with multi-keyframe direction and motion transfer
GPT-5.6 Terra
Jul 9, 2026OpenAI's balanced GPT-5.6 model for intelligence and cost
GPT-5.6 Luna
Jul 9, 2026OpenAI's cost-optimized GPT-5.6 model for high-volume workloads
TRELLIS 2
Jul 7, 2026Microsoft Research's open image-to-3D model
Microsoft MAI Models (Build 2026)
Jul 7, 2026Microsoft's unified AI model family from Build 2026
NVIDIA Cosmos 3
Jul 7, 2026Open physical-AI omnimodel for robotics and AV
Qwen-Audio-3.0-TTS
Jul 21, 2026Hosted Alibaba TTS across 16 languages, in a fast tier and a fidelity tier
Gemini 3.6 Flash
Jul 21, 2026Google's faster, sharper agentic-coding upgrade to 3.5 Flash
Qwen 3.8-Max
Jul 19, 2026Alibaba's 2.4-trillion-parameter flagship, currently in preview
June 2026 (15)
Claude Sonnet 5
Jun 30, 2026Anthropic's cheaper, near-Opus everyday model
Luma Uni-1.1
Jun 15, 2026Multimodal reasoning model that generates brand-consistent images and edits
Kimi K2.7 Code Highspeed
Jun 12, 2026Faster inference variant of Kimi's coding specialist
Claude Mythos 5
Jun 9, 2026Limited-availability Mythos-class model without Fable 5 safety classifiers
NVIDIA Nemotron 3 Nano
Jun 4, 2026Compact 30B open-weight model with configurable reasoning for agents
NVIDIA Nemotron 3 Super
Jun 4, 2026120B open-weight hybrid MoE for efficient multi-agent reasoning
NVIDIA Nemotron 3.5 Content Safety
Jun 4, 2026Multimodal 4B safety model for text and image moderation
NVIDIA GR00T N1.5 VLA
Jun 4, 2026Open foundation model for humanoid robot reasoning and control
GLM-5-Turbo
Jun 1, 2026Optimized GLM-5 variant for fast sequential task execution
GLM-5V-Turbo
Jun 1, 2026Multimodal coding and visual-reasoning agent model
MiniMax Speech 2.8 HD
Jun 1, 2026Ultra-realistic multilingual text-to-speech with sound tags
MiniMax Speech 2.8 Turbo
Jun 1, 2026Fast multilingual text-to-speech with natural flow
MiniMax Music 3.0
Jun 1, 2026Music generation with humanized vocals and elevated sound
NVIDIA Cosmos 3 Nano
Jun 1, 2026Efficient 8B physical-AI omni-model for workstations
NVIDIA Cosmos 3 Edge
Jun 1, 20264B physical-AI omni-model for real-time edge robotics
May 2026 (32)
GPT-Image-2
May 10, 2026OpenAI's latest image generation model
ComfyUI
May 19, 2026The node graph the rest of the field is measured against
Claude Opus 4.8
May 28, 2026Anthropic's powerful enterprise model from May 2026
GPT-5.5 Instant
May 5, 2026OpenAI's fast default ChatGPT model from May 2026
Cursor Composer 2.5
May 18, 2026Cursor's agentic coding model for multi-file software engineering
Gemini 3.5 Flash
May 19, 2026Google's fast, capable multimodal model from I/O 2026
Grok Build
May 14, 2026xAI's agentic coding CLI for autonomous software engineering
Gemini Omni
May 19, 2026Google's unified multimodal generation model
Runway Aleph 2.0
May 21, 2026In-context video editing model and Edit Studio
Gemini Spark
May 19, 2026Google's personal AI agent for proactive assistance
Mistral Vibe
May 28, 2026Mistral's unified work and coding agent
DeepSeek V4-Pro
May 31, 2026DeepSeek's open-weight model with permanent pricing
FLUX.2 [max]
May 15, 2026Black Forest Labs' top-tier image generation model
MiniMax M3
May 31, 2026MiniMax's 1M-context agentic frontier model
StepFun Step 3.7 Flash
May 29, 2026StepFun's 198B MoE vision-language model
Recraft V4
May 20, 2026Design and brand image generation with vector support
ElevenLabs Music v2
May 26, 2026AI music generation with professional controls
Nano Banana 2
May 22, 2026Google's fast text-to-image model via Fal
ElevenLabs Dubbing v2
May 28, 2026AI-powered video dubbing in multiple languages
Mistral Medium 3.5
May 22, 2026Mistral's mid-tier workhorse for reasoning, coding, and instruction
Udio v4
May 15, 2026AI music generation with stems and inpainting
Recraft V4.1
May 14, 2026Recraft's most advanced image model with photorealistic, vector, and utility variants
Topaz Astra
May 7, 2026Cloud AI video enhancement up to 4K
HappyHorse 1.0
May 3, 2026Alibaba flagship video with joint audio and multilingual lip-sync
NVIDIA Nemotron 3 Nano Omni
May 3, 2026One multimodal model for text, vision, audio, and video reasoning
Meshy 6
May 3, 2026Multi-image to production-grade 3D on next-gen Meshy
Decart Lucy 2.1 VTON
May 3, 2026Real-time virtual try-on in video
MiniMax M2.7
May 1, 2026Recursive self-improvement language model for real-world engineering
MiniMax M2.7 Highspeed
May 1, 2026Same M2.7 performance with significantly faster inference
Mistral Small 4
May 1, 2026Unified open-source small model for chat, reasoning, vision, and coding
Tripo v3.1
May 25, 2026Fast text-to-3D and image-to-3D generation
Rodin Gen-2
May 27, 2026Hyper3D's image-to-3D generation model
April 2026 (12)
Topaz Bloom
Apr 28, 2026Creative upscaling that adds realism to AI-generated images
Topaz Mobile
Apr 28, 2026Topaz image enhancement on iPhone
DeepSeek V4-Flash
Apr 24, 2026High-volume DeepSeek inference with a 1M-token context window
Hunyuan Hy3 Preview
Apr 22, 2026Tencent's latest open-source MoE flagship with tool use
Kimi K2.6
Apr 21, 2026Moonshot's open-weight multimodal successor with long-context coding stability
Claude Opus 4.7
Apr 16, 2026Frontier Opus model with higher-resolution vision and xhigh effort
Mistral Large 3
Apr 15, 2026Mistral's flagship open-weight multimodal frontier model
Ministral 3
Apr 15, 2026Mistral's edge family of small, dense open-source models
GLM-5.1
Apr 7, 2026MIT-licensed MoE flagship for 8-hour autonomous coding sessions
GLM-4.7-Flash
Apr 1, 2026Free universal GLM model with a 200K context window
GLM-4V-Flash
Apr 1, 2026Free vision model for image understanding and document snapshots
Grok 4.3
Apr 1, 2026xAI's long-context flagship with a 1M-token window
February 2026 (105)
Seedance 2.0
Feb 12, 2026ByteDance's Next-Gen Video Model with Native Audio-Video Joint Generation
NotebookLM
Feb 4, 2026Google's AI Research Assistant: The Ultimate Study Tool
Kling AI 3.0
Feb 6, 2026The frontier of cinematic video synthesis
Veo 3.1
Feb 5, 2026Google's state-of-the-art video generation model
Grok
Feb 5, 2026xAI's real-time AI assistant
DeepSeek
Feb 5, 2026The Efficiency Revolution: Frontier Intelligence at 1/100th the Cost
Suno
Feb 5, 2026Text-to-music & vocals with fast iteration
Llama
Feb 5, 2026Meta's open-source large language model
Mistral AI
Feb 5, 2026European open-source and commercial LLM
ElevenLabs
Feb 5, 2026High-quality TTS and voice tools
Cohere
Feb 5, 2026Enterprise-focused LLM platform
Qwen
Feb 5, 2026Alibaba's multilingual open-source LLM
Microsoft Phi
Feb 5, 2026Microsoft's efficient small language models
Gemma
Feb 5, 2026Google's open-source lightweight LLM
DBRX
Feb 5, 2026Databricks' high-performance open-source LLM
Resemble AI
Feb 5, 2026Voice generation and cloning tools
Sora 2
Feb 5, 2026OpenAI's state-of-the-art video model with audio
Kling 2.6 Pro
Feb 5, 2026Top-tier image-to-video with native audio generation
Kling AI
Feb 5, 2026Text/image-to-video generation (availability varies)
Runway
Feb 5, 2026Text/image-to-video creation suite with editing tools
Pika
Feb 5, 2026Text/image-to-video with Pikaffects (squish, melt, explode)
HeyGen
Feb 5, 2026Avatar and talking-head video generation
Ray2 Flash
Feb 5, 2026Fast video generation from Luma Dream Machine
Synthesia
Feb 5, 2026AI avatar video creation for teams
Hailuo 2.3 Fast
Feb 5, 2026Fast 1080p image-to-video from MiniMax
Claude Opus 4.6
Feb 6, 2026The ceiling of enterprise autonomy with 1M context
Midjourney
Feb 5, 2026High-end image generation with strong aesthetics
OmniHuman v1.5
Feb 5, 2026Audio-driven human animation from ByteDance
D-ID
Feb 5, 2026Talking avatar videos from images and scripts
Wan 2.1
Feb 5, 2026Open-source image-to-video with LoRA support
Hunyuan Video
Feb 5, 2026Tencent's high-quality open video model
Ideogram
Feb 5, 2026Text-to-image with strong typography (varies by model)
Leonardo AI
Feb 5, 2026Image generation with workflows and models
Wan 2.6 Text-to-Video
Feb 5, 2026Latest Wan model for text-to-video generation
Kaiber
Feb 5, 2026Stylized image/video animation for creators
Hunyuan Video 1.5
Feb 5, 2026Tencent's latest text-to-video model
Adobe Firefly
Feb 5, 2026Generative image tools inside Adobe ecosystem
LTX-2
Feb 5, 2026Fast text-to-video with audio support
Hunyuan 3D
Feb 5, 2026Tencent's high-quality 3D generation engine
Krea
Feb 5, 2026Creative image workflows (and some video features)
PixVerse
Feb 5, 2026Text/image-to-video with effects, transitions & swaps
Vidu Q2
Feb 5, 2026Shengshu's advanced image-to-video with better control
Viggle
Feb 5, 2026Character motion and meme-style video creation
GPT-Image 1.5
Feb 5, 2026OpenAI's high-fidelity image generation
GLM-5
Feb 11, 2026744B-parameter open-weight MoE flagship for agentic planning and execution
Qwen3-Coder-Next
Feb 6, 202680B parameter open-weight coding powerhouse
GPT-5.3 Codex
Feb 6, 2026The frontier model for complex reasoning and software architecture
Claude 4.6 Sonnet
Feb 6, 2026The industry standard for coding and nuanced instruction following
Runway Gen-4.5
Feb 6, 2026The industry standard for cinematic AI video generation
Gemini 3 Ultra
Feb 5, 2026Native multimodal intelligence with a 10M context window
Perplexity AI
Feb 5, 2026The conversational search engine that replaced traditional search
Consensus
Feb 5, 2026AI search engine for peer-reviewed scientific research
Tripo AI v3
Feb 5, 2026Instant high-quality 3D modeling from text and images
Luma Genie
Feb 5, 2026High-fidelity 3D asset generation from Luma Labs
Luma Dream Machine v2
Feb 5, 2026High-speed, high-realism video generation
FLUX.1 [pro]
Feb 5, 2026The new gold standard for prompt adherence and text rendering
Pika 2.0
Feb 5, 2026The creative suite for physics-defying video effects
Meshy AI v3
Feb 5, 2026Production-ready 3D assets in under 60 seconds
SAM3D v2
Feb 5, 2026Meta's Segment Anything 3D for high-fidelity reconstruction
Meshy AI
Feb 5, 2026Generate and refine 3D assets from text or images
Flux 2 Flex
Feb 5, 2026Fine-tuned control with adjustable inference
Grok 4.20
Feb 1, 2026xAI's 2M-context beta model with multi-agent capabilities
Recraft
Feb 5, 2026Design-forward image generation (logos, vectors, assets)
Flux Kontext
Feb 5, 2026Context-aware image generation and editing
Stable Diffusion 3.5
Feb 5, 2026Open-source image generation with flexibility
Magnific
Feb 5, 2026AI upscaling and enhancement for images
Wan 2.6 Image-to-Image
Feb 5, 2026Latest Wan for image variations and editing
Black Forest Labs
Feb 5, 2026FLUX image model family (provider site)
BRIA Eraser
Feb 5, 2026High-fidelity object removal from images
Microsoft TRELLIS
Feb 5, 2026Microsoft's advanced 3D generation from text or images
BRIA Video Eraser
Feb 5, 2026Object removal from video with high fidelity
Kaedim
Feb 5, 20262D-to-3D conversion for game assets
LightX Recamera
Feb 5, 2026Relight and recamera videos
Runway Gen-3 Alpha
Feb 5, 2026Advanced video editing and effects
MiniMax Music 2.0
Feb 5, 2026Advanced AI music generation with high-quality compositions
Stable Audio 2.5
Feb 5, 2026High-quality music and sound effects generation
Luma AI
Feb 5, 20263D capture + creative tools (incl. 3D/Video features)
ElevenLabs TTS Eleven-v3
Feb 5, 2026Multilingual text-to-speech with natural voice synthesis
Stable Diffusion
Feb 5, 2026Open image generation ecosystem (model + tools)
Lyria 2
Feb 5, 2026Google's latest music generation model
Canva
Feb 5, 2026Design suite with built-in AI generation features
Sonauto v2.2
Feb 5, 2026CD-quality music with superior vocals
Descript
Feb 5, 2026Audio/video editing with AI features
ElevenLabs Sound Effects v2
Feb 5, 2026Advanced sound effects generation
Flux 1 [schnell]
Feb 5, 2026Fast Flux variant for rapid image generation
Imagen 3
Feb 5, 2026Google's high-quality text-to-image model
Recraft V3
Feb 5, 2026Vector art and brand-style image generation
Ideogram V3
Feb 5, 2026Exceptional typography and text rendering
Topaz Photo AI
Feb 5, 2026Image enhancement (denoise/sharpen/upscale)
Flux 1 [dev]
Feb 5, 2026Development Flux for advanced control
Spline
Feb 5, 20263D design tool (with AI features depending on product)
Ovis Image
Feb 5, 2026Quick text rendering for marketing graphics
LongCat Image
Feb 5, 2026Multilingual text rendering and photorealism
Bagel
Feb 5, 20267B multimodal model for text and images
Flux Realism LoRA
Feb 5, 2026Photorealistic Flux with LoRA fine-tuning
Flux LoRA
Feb 5, 2026Customizable Flux with LoRA fine-tuning
MiniMax TTS
Feb 5, 2026Multilingual text-to-speech with streaming
Shap-E
Feb 5, 2026OpenAI's conditional 3D model generation
Point-E
Feb 5, 2026OpenAI's fast point cloud generation
DreamFusion
Feb 5, 2026Text-to-3D via NeRF with score distillation
Get3D
Feb 5, 2026NVIDIA's high-quality 3D mesh generation
Topaz Video Enhance AI
Feb 5, 2026Professional video upscaling and enhancement
CapCut
Feb 5, 2026AI-powered video editing with enhancement features
Zero-1-to-3
Feb 5, 2026View-consistent image-to-3D generation
Instant3D
Feb 5, 2026Fast single-image 3D generation
January 2026 (15)
Kimi k1.5
Jan 31, 2026The 'Next DeepSeek' Movement: o1-Level Reasoning at 1/100th the Cost
Qwen 2.5-VL
Jan 31, 2026The Open Vision-Reasoner: SOTA Multimodal Performance
Llama 3.2 Vision
Jan 31, 2026Meta's Open Multimodal Standard
Pixtral Large
Jan 31, 2026The Open Vision Frontier: 124B Multimodal Power
InternVL 2.5
Jan 31, 2026The Open-Source Vision Giant: 78B Multimodal Leader
Flora
Jan 31, 2026The Workflow Canvas: Figma for Generative AI
Z-Image
Jan 1, 2026Ultra-fast photorealistic image generation with bilingual text rendering
Qwen-Image
Jan 1, 2026Open-source 20B model with commercial-grade text rendering and advanced image editing
FLUX.2 Pro
Jan 1, 2026The Open Image Standard: The Midjourney Killer
Kimi K2.5
Jan 27, 2026Moonshot's open-weight multimodal generalist with agent swarms
Hunyuan TurboS
Jan 10, 2026Tencent's fast, cost-efficient flagship Hunyuan model
Baidu ERNIE 4.5
Jan 1, 2026Open-source MoE LLM with strong Chinese NLP and multimodal capabilities
GLM-4.5
Jan 1, 2026Advanced multilingual LLM with enhanced reasoning and long-context support
Hymotion 1.0
Jan 1, 2026Open-source text-to-3D motion model with 200+ motion categories and production-ready exports
Manus AI
Jan 1, 2026Autonomous AI agent for complex multi-step workflows and research automation
December 2025 (5)
Seedance 1.5 pro
Dec 16, 2025Cinematic audio-video joint generation with lip-sync and dialect support
DeepSeek V3.2
Dec 1, 2025The 128K-context MoE flagship that introduced sparse attention
GLM-4.7
Dec 1, 2025Strong general-reasoning model with interleaved thinking
GLM-4.6V
Dec 1, 2025Vision-language model for visual reasoning and UI replication
Tripo 4.0
Dec 1, 2025Tripo's latest high-fidelity 3D generation model
November 2025 (5)
Claude Opus 4.5
Nov 24, 2025First Claude model with the effort parameter and context compaction
Hunyuan A13B Instruct
Nov 15, 2025Tencent's efficient small-scale MoE instruct model
BRIA FIBO Lite
Nov 11, 2025Fast, lightweight FIBO pipeline designed for speed, efficiency, and on-prem deployment
Cursor Composer 1
Nov 1, 2025Cursor's first-generation agentic coding model
Grok 4.1 Fast
Nov 1, 2025xAI's high-volume, 2M-context workhorse model
October 2025 (2)
September 2025 (5)
Claude Sonnet 4.5
Sep 29, 2025Balanced Sonnet model with major coding and agentic improvements
HunyuanImage 3.0
Sep 28, 2025Tencent's 80B-parameter open-source MoE image generator
Hunyuan 2.0 Think
Sep 15, 2025The deep-thinking variant of Hunyuan 2.0
GLM-4.6
Sep 1, 2025Mid-range coding and tool-calling model with 200K context
Tripo 3.5
Sep 1, 2025Mid-generation upgrade between Tripo 3 and 4
August 2025 (1)
July 2025 (4)
HunyuanWorld
Jul 26, 2025Tencent's open-source immersive 3D world generator
BRIA RMBG 2.0
Jul 24, 2025High-accuracy background removal model trained on a licensed, professionally labeled dataset
GLM-4.5-Air
Jul 1, 2025Cost-efficient reasoning, coding, and agent model
Voxtral
Jul 1, 2025Mistral's open-weight speech understanding and TTS models
June 2025 (10)
BRIA FIBO
Jun 26, 2025Open-source JSON-native text-to-image model built for controllable, enterprise-safe generation
Hunyuan 2.0 Instruct
Jun 20, 2025Tencent's general-purpose instruction-tuned Hunyuan 2.0 model
Hunyuan3D 2.1
Jun 13, 2025Tencent's latest open-source high-fidelity 3D asset generator
Seedance 1.0
Jun 11, 2025Fast, inference-efficient 1080p video with native multi-shot storytelling
FLUX.2 [schnell]
Jun 1, 2025Fast local FLUX.2 generation for personal hardware
FLUX.2 [dev]
Jun 1, 2025Open-weight FLUX.2 for research and commercial use
ElevenLabs Text to Voice v3
Jun 1, 2025Generate custom synthetic voices from text descriptions
NVIDIA Nemotron Parse
Jun 1, 2025Layout-aware document parsing that goes beyond OCR
Pika 2.2
Jun 1, 2025Pika's refined model with stronger realism and camera control
Topaz Image Web
Jun 1, 2025Browser-based AI image enhancement workflows
April 2025 (5)
Qwen 3
Apr 28, 2025Alibaba's open-source MoE flagship with thinking modes
Llama 4 Maverick
Apr 5, 2025Meta's open-weight flagship with native multimodal reasoning
Llama 4 Scout
Apr 5, 2025Long-context, efficient open multimodal model for edge and single-GPU use
Kling 3.0 Master
Apr 1, 2025Premium tier of Kling 3.0 with best quality
Runway Gen-4
Apr 1, 2025Runway's next-generation model for consistent characters and camera
March 2025 (5)
Gemini 2.5 Pro
Mar 25, 2025Google's high-performance reasoning model with advanced coding
Hunyuan T1
Mar 21, 2025Tencent's Mamba-powered deep-thinking reasoning model
Gemma 3
Mar 12, 2025Google's open multimodal model for research and developers
Wan 2.0
Mar 1, 2025Earlier open-source Wan video generation model
Kling Image 2.0
Mar 1, 2025Kling's image generation model with style control
February 2025 (1)
January 2025 (4)
Qwen 2.5-Max
Jan 28, 2025Alibaba's closed-API flagship before Qwen 3
Perplexity Sonar Reasoning Pro
Jan 21, 2025Chain-of-thought reasoning model for multi-step logical analysis
DeepSeek R1
Jan 20, 2025The open-weight reasoning model that sparked the efficiency revolution
Kling 2.5
Jan 1, 2025Improved physics and expressive movement in Kling video
December 2024 (9)
Veo 2
Dec 16, 2024Google's high-quality 1080p video generation model
Gemini 2.0 Flash
Dec 11, 2024Google's low-latency agentic model with native tool use
Llama 3.3
Dec 6, 2024Efficient 70B open model matching 405B quality
ElevenLabs Flash v2.5
Dec 1, 2024Ultra-low-latency text-to-speech for real-time voice agents
Llama-3.1-Nemotron Ultra
Dec 1, 2024NVIDIA-aligned 253B Llama 3.1 for helpfulness and instruction following
Llama-3.1-Nemotron Super
Dec 1, 2024NVIDIA-aligned 49B Llama 3.1 for balanced performance
Llama-3.1-Nemotron Nano
Dec 1, 2024NVIDIA-aligned 8B Llama 3.1 for efficient inference
TRELLIS Mini
Dec 1, 2024Compact open-source image-to-3D model from Microsoft
TRELLIS Large
Dec 1, 2024High-quality open-source image-to-3D from Microsoft
November 2024 (4)
Qwen 2.5-Coder
Nov 12, 2024Alibaba's open coding-specialist model
Hunyuan Large
Nov 4, 2024Tencent's 389B-parameter open-source MoE language model
Perplexity Sonar Pro
Nov 1, 2024Advanced search model with deeper reasoning and richer citations
Runway Act-One
Nov 1, 2024Performance-driven character animation from video
October 2024 (8)
Stable Diffusion 3.5 Large
Oct 22, 2024Stability AI's largest 3.5 model with best quality
FLUX.1.1 [pro]
Oct 2, 2024Ultra-realistic FLUX.1 update with faster generation
FLUX.1 Fill [pro]
Oct 2, 2024Advanced inpainting and outpainting FLUX model
FLUX.1 Canny
Oct 2, 2024Canny-edge-guided image generation and editing
FLUX.1 Depth
Oct 2, 2024Depth-map-guided image generation and editing
Kling 2.0
Oct 1, 2024Kling's standard model for cinematic video
Pika 1.5
Oct 1, 2024Pika's upgrade with improved motion and effects
Runway Frames
Oct 1, 2024Image generation model with strong style control
September 2024 (1)
August 2024 (2)
July 2024 (2)
June 2024 (6)
Stable Diffusion 3
Jun 12, 2024Stability AI's first multimodal-diffusion Transformer image model
Stable Diffusion 3 Medium
Jun 12, 2024Efficient SD3 variant for consumer hardware
Kling 1.5
Jun 6, 2024Kling's first widely available video generation model
ElevenLabs Multilingual Speech to Speech v2
Jun 1, 2024Real-time multilingual voice conversion that preserves emotion and content
Topaz Gigapixel
Jun 1, 2024AI-powered image upscaling up to 8x with detail recovery
Tripo 2.0
Jun 1, 2024Earlier generation of Tripo's text- and image-to-3D pipeline