TAG • CURATED

API AI Tools

AI tools with API access enable programmatic integration into your applications and workflows. Build custom solutions, automate content generation, and scale production with direct API access to AI capabilities.

RESULTS
255 tools • curated
API platform for 600+ generative AI models
Added Feb 5, 2026
Cloud-based serverless GPU platform providing unified API access to over 600 generative AI models across multiple modalities including image generation, video generation, audio synthesis, 3D creation,...
Why: Largest collection of generative AI models accessible via unified API, making it the most comprehensive platform for multi-modal AI development.
Enterprise Best for Multi-Model Access Visit
The LLM-Ready Web Scraper: Turn Websites into Markdown
Added Jan 31, 2026
Firecrawl is the industry-standard tool for turning entire websites into clean, LLM-ready markdown
Why: Firecrawl is the leader of the 'LLM-Data' movement. We picked it because it's the first scraper that actually understands what AI models need: clean, noise-free markdown without the overhead of traditional scraping libraries.
Enterprise Best for AI Data Extraction Visit
The Post-Search Era: The End of the Blue Link
Added Jan 1, 2026
Perplexity Comet is the spearhead of the 'Post-Search' era, a fundamental shift from ad-driven link lists to source-driven answers
Why: Perplexity Comet represents the death of the traditional search engine. We picked it because it's the first agentic browser to prove that autonomous web navigation and source-backed reasoning are more valuable than a list of 'blue links.'
Free Best for Post-Search Research Visit
The Native Agentic Layer: The Browser as an OS
Added Jan 31, 2026
Google Chrome has evolved from a simple browser into a native agentic layer powered by Gemini 3
Why: We added Chrome to the agentic category because it represents the first time a mainstream browser has integrated a native reasoning engine that can autonomously navigate the web on behalf of the user.
Freemium Best for Native Web Automation Visit
The Invisible OS: Pure Execution via Messaging
Added Jan 27, 2026
Moltbot (also known as Clawdbot) is the spearhead of the 'Invisible OS' movement, a shift away from fragmented apps and toward pure, autonomous execution via messaging
Why: Moltbot represents the death of the 'app for everything' era. We picked it because it's the first agentic assistant to prove that reasoning-based execution through simple chat is more powerful than manual task management in 10+ different apps.
Freemium Best for Agentic Automation Visit
OpenAI's latest image generation model
Added May 10, 2026
GPT-Image-2 is OpenAI's image generation model, first announced on April 21, 2026, and available through the API in early May 2026
Why: GPT-Image-2 is OpenAI's most capable image model to date, with notably better text-in-image accuracy. It is a natural choice for OpenAI API users who want image generation alongside text and audio in a single platform.
Freemium Best for OpenAI Image API Visit
30-second 4K video with native audio and up to 50 reference inputs
New this month Added Aug 4, 2026
Seedance 2
Why: The longest single-run generation of any current video model at 4K, and the 50-reference input system is the most direct answer yet to character consistency, the problem that breaks most AI video work. Availability is the constraint: it ships inside ByteDance's own apps first, and the previous generation's international rollout was postponed indefinitely.
Freemium Best for Long Clips Visit
The management layer for AI agent workforces
Added Feb 6, 2026
A new enterprise platform designed to deploy, manage, and oversee AI agents as if they were human employees
Why: OpenAI Frontier is like a 'Manager for Robots.' Instead of you having to talk to 10 different AI tools one by one, Frontier lets you manage them all like a team of employees. It makes sure they stay safe, follow the rules, and work together to get big jobs done for your business.
Enterprise Best for Agent Management Visit
ByteDance's Next-Gen Video Model with Native Audio-Video Joint Generation
Added Feb 12, 2026
Seedance 2
Why: Seedance 2.0 represents the new frontier of multimodal generation, being one of the first models to generate high-fidelity audio and video simultaneously with extreme temporal consistency and physics-based realism.
Best for Cinematic Visit
The Open-Source Scraping Engine: High-Performance LLM Crawling
Added Jan 31, 2026
Crawl4AI is an open-source, high-performance web crawling and scraping engine specifically optimized for large language models
Why: Crawl4AI is the leading open-source alternative to proprietary scraping APIs. We picked it because it offers the most powerful 'local-first' crawling experience, giving developers full control over their data extraction pipeline without the per-page costs of cloud services.
Free Best for Open-Source Crawling Visit
Platform for prototyping with Google's Gemini models
Added Feb 5, 2026
Web-based integrated development environment for prototyping and building applications with Google's generative AI models
Why: Official Google platform providing direct access to Gemini models with excellent developer tools and seamless API integration.
Freemium Best for Gemini Models Visit
Anthropic's frontier model, currently first on the Artificial Analysis Intelligence Index
New this month Added Aug 4, 2026
Claude Opus 5 is Anthropic's flagship model, released 24 July 2026 with a 1M-token context window and five selectable effort levels (low, medium, high, xhigh, max)
Why: It is the current number one on the independent Artificial Analysis Intelligence Index, and it got there while costing less per task than the model it displaced: $2.03 average per index task against Fable 5's $2.75. The effort dial is the reason to pick it over a fixed-tier model, because one integration covers cheap high-volume calls and expensive long-horizon agent runs.
Freemium Best Frontier Model Overall Visit
The frontier of cinematic video synthesis
Added Feb 6, 2026
Kling AI is a state-of-the-art video generation platform capable of producing high-fidelity cinematic content
Why: Kling AI is the current king of AI movies. It can create high-quality video clips that are 2 minutes long, which is like an eternity in AI time, while keeping the characters and the physics (like how water splashes or hair moves) looking perfectly real. It's the first tool that lets professional filmmakers create a whole scene without the video 'glitching' halfway through.
Freemium Best for Filmmaking Visit
Unified API for multiple LLM models
Added Feb 5, 2026
Unified API platform providing access to multiple large language models from different providers through a single API interface
Why: Best unified API for accessing multiple LLM providers, making it easy to switch models or use multiple models in one application.
Enterprise Best for Model Flexibility Visit
Google's state-of-the-art video generation model
Added Feb 5, 2026
Generates high-quality videos from text prompts or images using Google DeepMind's Veo 3
Why: Google's state-of-the-art video model with top-tier cinematic quality and flexible input options including reference and frame control.
Paid Best for Cinematic Visit
AI system that translates natural language into code
Added Feb 5, 2026
AI system developed by OpenAI that translates natural language prompts into code across multiple programming languages
Why: Foundation technology powering GitHub Copilot and enabling natural language to code translation.
Paid Best for Code Generation Visit
API access to thousands of models on Hugging Face
Added Feb 5, 2026
Provides API access to thousands of machine learning models hosted on the Hugging Face Hub
Why: Largest model repository with API access, making it the go-to platform for accessing diverse AI models.
Enterprise Best for Model Variety Visit
Fast inference platform for AI models
Added Feb 5, 2026
High-performance inference platform providing ultra-fast API access to large language models and other AI models
Why: Fastest inference platform available, making it ideal for real-time applications requiring low latency.
Enterprise Best for Speed Visit
Moonshot AI's 2.8-trillion-parameter open-weight flagship
Added Jul 16, 2026
Kimi K3 is Moonshot AI's open-weight model released July 16, 2026, built at roughly 2
Why: Kimi K3 is one of the most credible open-weight challengers to closed frontier models this year, aggressive enough on pricing and scale that it moved markets, Fortune covered it as a 'DeepSeek shock' moment for AI stocks.
Freemium Best for Open-Weight Frontier Performance Visit
Design platform with multiple AI tools and licensed content
Added Feb 5, 2026
Graphic design platform offering multiple AI-powered tools including F Lite image generator (trained on licensed data), image editing, video generation, icon generation, AI image classification, and a...
Why: Unique combination of AI tools and licensed content, ensuring commercial compliance for design projects.
Freemium Best for Licensed Content Visit
xAI's real-time AI assistant
Added Feb 5, 2026
Grok is xAI's AI assistant integrated into X (formerly Twitter) with real-time access to platform data and a more conversational, edgy tone
Why: xAI's AI assistant with unique real-time X platform integration and distinctive conversational style for social media context.
Paid Best for Real-time Visit
The Efficiency Revolution: Frontier Intelligence at 1/100th the Cost
Added Feb 5, 2026
DeepSeek is the architect of the 'DeepSeek movement,' a fundamental shift in AI development that prioritizes extreme efficiency over raw compute
Why: DeepSeek changed the game by proving that 'expensive' doesn't always mean 'better.' We picked it because it's the first model family to offer true frontier-level reasoning (R1), general intelligence (V3), and advanced vision/OCR (VL2) with an open-weight philosophy and an API price point that makes proprietary models look obsolete.
Freemium Best for Cost-Efficiency Visit
Meta's open-source large language model
Added Feb 5, 2026
Llama is Meta AI's open-source large language model family with multiple versions: Llama (February 2023), Llama 2 (July 2023), Llama 3 (April 2024), Llama 3
Why: Meta's flagship open-source LLM with strong performance, extensive model sizes, and permissive licensing for research and commercial use.
Free Best for Open Source Visit
Cloud-based online IDE for web development
Added Feb 5, 2026
Cloud-based online IDE focused on web application development
Why: Best cloud IDE for web development with instant setup and collaboration.
Freemium Best for Web Development Visit
European open-source and commercial LLM
Added Feb 5, 2026
Mistral AI provides high-performance large language models with both open-source and commercial offerings
Why: European LLM provider with strong open-source offerings, multilingual capabilities, and focus on data privacy and compliance.
Freemium Best for Europe Visit
High-quality TTS and voice tools
Added Feb 5, 2026
Generates realistic text-to-speech voiceovers with natural intonation and emotion
Why: Best voice quality combined with reliable API for production pipelines requiring consistent, natural-sounding narration.
Freemium Best for Narration Visit
Enterprise-focused LLM platform
Added Feb 5, 2026
Cohere provides enterprise-grade large language models including Command, Command R, Command R+, Command R7
Why: Enterprise-focused LLM with strong RAG capabilities, multilingual support, and emphasis on accuracy and safety for business use cases.
Enterprise Best for Enterprise Visit
Omni-modal video with native stereo audio, at 2K
New this month Added Jul 31, 2026
MiniMax H3, released 31 July 2026, reads text, images, video and audio as one shared context and returns video with native stereo sound at up to 2K and 15 seconds
Why: The native stereo audio is the part that matters — most 2K video models still hand you a silent clip and leave you to score it. Generating sound in the same pass as the picture removes a whole step, and MiniMax prices the 2K tier well under the mainstream alternatives.
Paid Best for Video With Sound Visit
Alibaba's multilingual open-source LLM
Added Feb 5, 2026
Qwen is Alibaba Cloud's family of large language models with multiple versions: Qwen-1
Why: Alibaba's high-performance multilingual LLM with strong Chinese language support, cost-efficient pricing, and comprehensive open-source availability.
Freemium Best for Multilingual Visit
Microsoft's efficient small language models
Added Feb 5, 2026
Microsoft Phi is a family of small, efficient language models designed for high performance with minimal parameters
Why: Microsoft's efficient small language models with strong reasoning capabilities, MIT licensing, and optimized for resource-constrained environments.
Free Best for Efficiency Visit
Google's open-source lightweight LLM
Added Feb 5, 2026
Gemma is Google DeepMind's family of open-source large language models, serving as lightweight versions of Gemini
Why: Google's open-source LLM family with strong performance, permissive licensing, and specialized variants for vision and medical applications.
Free Best for Research Visit
The 'Next DeepSeek' Movement: o1-Level Reasoning at 1/100th the Cost
Added Jan 31, 2026
Kimi k1
Why: Kimi k1.5 is the first model to prove that o1-level reasoning is achievable through efficient, open-weight architectures. We selected it because it consistently matches or exceeds Claude 4.5 in technical benchmarks (AIME, MATH-500) while offering a 2M context window and a significantly lower API price point, making frontier intelligence accessible to everyone.
Freemium Best for Technical Reasoning Visit
The Open Vision-Reasoner: SOTA Multimodal Performance
Added Jan 31, 2026
Qwen 2
Why: We added Qwen 2.5-VL to the Open Frontier movement because it is currently the highest-performing open-weight vision model. It proves that open source can lead in multimodal reasoning, especially for tasks requiring high-resolution OCR and long-form video understanding.
Free Best for Open Vision Reasoning Visit
Databricks' high-performance open-source LLM
Added Feb 5, 2026
DBRX is a mixture-of-experts transformer model developed by Databricks and Mosaic ML
Why: Databricks' high-performance open-source LLM with strong benchmark results, efficient MoE architecture, and permissive licensing.
Enterprise Best for Performance Visit
Meta's Open Multimodal Standard
Added Jan 31, 2026
Llama 3
Why: We included Llama 3.2 Vision because it is the most widely supported open multimodal model in the world. Its integration into almost every AI tool and framework makes it the 'default' choice for open-weight vision reasoning.
Free Best for Open Ecosystem Support Visit
Voice generation and cloning tools
Added Feb 5, 2026
Creates synthetic voices and voiceovers from text with voice cloning capabilities
Why: Good option when you need voice tooling and APIs for production workflows requiring voice cloning and customization.
Best for Voice Visit
The Open Vision Frontier: 124B Multimodal Power
Added Jan 31, 2026
Pixtral Large is Mistral AI's flagship 124B parameter multimodal model, designed to compete directly with GPT-4o and Claude 3
Why: We added Pixtral Large because it represents the peak of European open-weight AI. It is one of the few open models that truly matches the visual reasoning depth of the top proprietary models, making it essential for the Open Frontier movement.
Freemium Best for Complex Visual Reasoning Visit
The Open-Source Vision Giant: 78B Multimodal Leader
Added Jan 31, 2026
InternVL 2
Why: We included InternVL 2.5 because it is a consistent leaderboard champion. It often outperforms much larger models in visual reasoning and OCR, making it a critical tool for developers who need GPT-4 level vision without the proprietary lock-in.
Free Best for Leaderboard-Topping Vision Visit
OpenAI's state-of-the-art video model with audio
Added Feb 5, 2026
Creates richly detailed, dynamic video clips with native audio generation from text prompts or images using OpenAI's Sora 2 model
Why: OpenAI's flagship video model with native audio generation, representing state-of-the-art quality in video synthesis.
Paid Best for Cinematic Visit
Top-tier image-to-video with native audio generation
Added Feb 5, 2026
Generates cinematic videos from images using Kling 2
Why: Best-in-class motion fluidity + native audio support, making it the top choice for cinematic image-to-video generation.
Paid Best for Cinematic Visit
Text/image-to-video generation (availability varies)
Added Feb 5, 2026
Generates videos from text prompts or images using Kling's video generation models
Why: Often strong motion and quality when available, with cinematic visuals and fluid motion capabilities.
Best for Video Visit
Text/image-to-video creation suite with editing tools
Added Feb 5, 2026
Generates videos from text or images and provides a complete web-based editing suite
Why: Best all-in-one product workflow combining video generation with professional editing tools in a single platform.
Paid Best for Workflow Visit
Avatar and talking-head video generation
Added Feb 5, 2026
Creates talking-head and AI avatar videos from text scripts with multilingual support
Why: Easy path to presenter-style videos for teams with multilingual support and professional avatar quality.
Enterprise Best for Video Visit
Fast video generation from Luma Dream Machine
Added Feb 5, 2026
Creates realistic visuals with natural, coherent motion using Luma's Ray2 Flash model optimized for speed
Why: Speed + quality balance for quick iterations with fast generation times and reliable motion quality.
Freemium Best for Speed Visit
AI avatar video creation for teams
Added Feb 5, 2026
Creates presenter-style videos from text scripts using AI avatars with professional quality
Why: One of the most established options for corporate training and explainers with proven enterprise reliability.
Enterprise Best for Avatars Visit
Fast 1080p image-to-video from MiniMax
Added Feb 5, 2026
Advanced fast image-to-video generation with up to 1080p resolution using MiniMax's Hailuo 2
Why: Speed + high resolution (1080p Pro tier) combination making it ideal for fast, high-quality video generation.
Paid Best for Speed Visit
The ceiling of enterprise autonomy with 1M context
Added Feb 6, 2026
Anthropic's most powerful model, designed for autonomous software engineering and complex reasoning
Why: Claude Opus 4.6 is like a super-smart digital architect. While most AI can only write short snippets, Opus can 'see' your entire project (up to 1 million words) at once. It doesn't just help you code; it can actually build complex software systems from scratch, making it the best choice for big companies that need an AI 'teammate' rather than just a chatbot.
Enterprise Best for Autonomy Visit
Audio-driven human animation from ByteDance
Added Feb 5, 2026
Generates video from image and audio input with correlated emotions and movements using ByteDance's OmniHuman v1
Why: Best for realistic talking avatars with emotional sync, providing the most natural audio-driven human animation available.
Paid Best for Avatar Visit
DeepSeek's open-weight model with permanent pricing
Added May 31, 2026
DeepSeek V4-Pro is a high-performance language model from DeepSeek
Why: DeepSeek V4-Pro stands out for combining frontier-level performance with transparent, permanent pricing and open weights. It is a practical choice for teams that want to self-host or avoid unpredictable API costs.
Freemium Best for Predictable Pricing Visit
Talking avatar videos from images and scripts
Added Feb 5, 2026
Animates a face image into talking-head video from text or audio input
Why: Fast route to talking-head content from a single image with reliable lip-sync and natural expressions.
Enterprise Best for Avatars Visit
Open-source image-to-video with LoRA support
Added Feb 5, 2026
Generates high-quality videos with motion diversity from images using Wan 2
Why: Open-source + LoRA customization for advanced users who need fine-tuned control and self-hosting capabilities.
Free Best for Open Source Visit
The Workflow Canvas: Figma for Generative AI
Added Jan 31, 2026
Flora is a collaborative AI design canvas that moves beyond the prompt box and into node-based workflow orchestration
Why: Flora is built for more than one person working on the same graph at the same time, which most node canvases are not. If the bottleneck in your work is handing a workflow to a colleague rather than the workflow itself, that is what it solves. ComfyUI gives you more control and Invoke gives you a better editing canvas, so pick this one for the collaboration.
Freemium Best for AI Design Workflows Visit
Tencent's high-quality open video model
Added Feb 5, 2026
High-quality image-to-video generation from Tencent using open-source Hunyuan Video models
Why: Strong open-source option with good quality, making it ideal for self-hosting and customization workflows.
Free Best for Open Source Visit
Text-to-image with strong typography (varies by model)
Added Feb 5, 2026
Generates images from text prompts with exceptional typography and text rendering capabilities
Why: Great for posters, logos, and brand mockups where accurate text rendering is critical.
Freemium Best for Images Visit
Image generation with workflows and models
Added Feb 5, 2026
Generates and edits images with a creator-friendly UI and extensive model library
Why: Good all-around image tool with comprehensive workflow features for concept art and production pipelines.
Freemium Best for Images Visit
Latest Wan model for text-to-video generation
Added Feb 5, 2026
Generates videos from text prompts using Wan 2
Why: Latest iteration of Wan with improved quality and control, representing the cutting edge of Wan's text-to-video capabilities.
Best for Video Visit
MiniMax's 1M-context agentic frontier model
Added May 31, 2026
MiniMax M3 is a 1-million-token-context agentic frontier model released on May 31, 2026
Why: MiniMax M3's 1M context window makes it competitive for tasks that require digesting entire codebases, books, or video transcripts in a single pass. It is a strong option for long-context agentic applications.
Freemium Best for 1M Context Visit
Tencent's latest text-to-video model
Added Feb 5, 2026
Generates videos from text prompts with high quality and motion control using Tencent's Hunyuan Video 1
Why: Tencent's flagship T2V model with strong performance, making it a top choice for high-quality text-to-video generation.
Best for Video Visit
Ultra-fast photorealistic image generation with bilingual text rendering
Added Jan 1, 2026
Generates high-quality photorealistic images from text prompts using Tongyi-MAI's Z-Image model with Single-Stream Diffusion Transformer (S3-DiT) architecture
Why: Ultra-fast photorealistic generation with superior bilingual text rendering, making it ideal for designs requiring text-in-image accuracy.
Freemium Best for Speed Visit
StepFun's 198B MoE vision-language model
Added May 29, 2026
StepFun Step 3
Why: Step 3.7 Flash offers a competitive Chinese-frontier multimodal model with an MoE architecture that balances capability and inference cost. It is a useful option for vision-language applications and for teams exploring alternatives to US models.
Freemium Best for Efficient VLM Visit
Fast text-to-video with audio support
Added Feb 5, 2026
Generates videos from text with native audio generation support using LTX-2 model
Why: Speed + audio in one model for complete video generation, eliminating the need for separate audio synthesis steps.
Best for Speed Visit
Open-source 20B model with commercial-grade text rendering and advanced image editing
Added Jan 1, 2026
Generates high-quality images from text prompts using Alibaba's Tongyi Qianwen 20-billion parameter MMDiT model
Why: Top-performing open-source model with exceptional text rendering and advanced image editing capabilities, optimized for efficient deployment.
Free Best for Text Rendering Visit
Tencent's high-quality 3D generation engine
Added Feb 5, 2026
Generates high-quality 3D models from text descriptions, images, or sketches using Tencent's Hunyuan 3D engine
Why: Tencent's comprehensive 3D generation engine with support for multiple input types and professional output formats, making it ideal for production workflows.
Best for 3D Assets Visit
The Open Image Standard: The Midjourney Killer
Added Jan 1, 2026
FLUX
Why: FLUX.2 represents the shift toward 'High-End Open Source.' We picked it because it matches Midjourney's aesthetic quality while offering the transparency and customizability that only an open-weight model can provide.
Freemium Best for Open-Weight Quality Visit
Text/image-to-video with effects, transitions & swaps
Added Feb 5, 2026
Generates short videos from text prompts or images with an extensive effects library, smooth transitions between scenes, and advanced object/person/background swapping capabilities
Why: Comprehensive effects library + seamless transitions + object swapping in one platform, making it ideal for creative video work requiring multiple transformation capabilities.
Freemium Best for Effects Visit
Shengshu's advanced image-to-video with better control
Added Feb 5, 2026
Generates high-quality videos from images using Shengshu's Vidu Q2 model with improved quality and control options compared to Q1
Why: Better quality and control compared to Q1, making it the preferred choice for high-quality image-to-video generation.
Best for Cinematic Visit
OpenAI's high-fidelity image generation
Added Feb 5, 2026
Generates high-fidelity images from text prompts using OpenAI's GPT-Image 1
Why: OpenAI's flagship image generation model with state-of-the-art prompt following and detail preservation, representing the cutting edge of text-to-image quality.
Paid Best for Quality Visit
80B parameter open-weight coding powerhouse
Added Feb 6, 2026
Alibaba's latest open-weight model specialized for coding
Why: Qwen3-Coder-Next is the best 'Private Brain' for coders. Most AI tools send your secret code to the internet, but this one can live entirely on your own computer. It's just as smart as the big paid tools, but it keeps your work 100% private and safe.
Free Best for Open Coding Visit
The platform for frontend and AI-first applications
Added Feb 5, 2026
Vercel is the default deployment platform for modern web apps
Why: The vertical integration of v0.dev and Edge compute makes Vercel the fastest path from prompt to production for AI applications. It's the only platform that optimizes the entire stack from generative UI to low-latency model inference at the edge, making it indispensable for high-performance AI startups.
Enterprise Best for Deployment Visit
The Bloomberg Terminal for AI agent observability
Added Feb 5, 2026
LangSmith provides full-stack observability for LLM applications
Why: LangSmith is like a 'Security Camera' for your AI. Sometimes AI gets confused or makes mistakes, and LangSmith lets you watch exactly what it was thinking so you can fix it. It's the best way to make sure your AI stays helpful and doesn't waste money.
Enterprise Best for Observability Visit
The managed vector database for long-term AI memory
Added Feb 5, 2026
Pinecone is a high-performance vector database designed for RAG (Retrieval-Augmented Generation)
Why: Pinecone is the AI's 'Infinite Filing Cabinet.' While most AI forgets what you said yesterday, Pinecone stores all your important info in a way the AI can find in a split second. It's what lets an AI 'remember' your specific business facts forever.
Enterprise Best for Memory Visit
The open-source Firebase alternative with Vector support
Added Feb 5, 2026
Supabase provides a unified backend stack including a Postgres database, authentication, and storage
Why: Supabase is the 'All-in-One Toolbox' for building AI apps. It gives you a database, a way for users to log in, and a place for the AI to store its memory all in one spot. It's the easiest way to go from an idea to a working app without needing 10 different services.
Enterprise Best for Backend Visit
Serverless GPU compute for heavy AI workloads
Added Feb 5, 2026
Modal allows developers to run Python code in the cloud with instant access to GPUs
Why: Modal is like 'Renting a Supercomputer' by the second. Usually, you need very expensive computers to train AI, but Modal lets you use theirs only when you need it. It's the cheapest and fastest way for small teams to do big AI work.
Enterprise Best for Compute Visit
The frontier model for complex reasoning and software architecture
Added Feb 6, 2026
OpenAI's most advanced model to date, featuring a 2M context window and specialized training for complex multi-step reasoning
Why: GPT-5.3 Codex is the 'World's Smartest Planner.' While other AI tools are good at chatting, this one is built for solving huge, difficult problems like planning how a whole software system should work. It has a massive memory (2 million words) so it never loses track of the big picture.
Paid Best for Reasoning Visit
The industry standard for coding and nuanced instruction following
Added Feb 6, 2026
Anthropic's flagship model, optimized for high-speed coding and perfect adherence to complex XML-based system prompts
Why: Claude 4.6 Sonnet is the 'Perfect Student' for following directions. It is famous for doing exactly what you ask without getting confused. It also has a special 'Computer Use' feature where it can actually move the mouse and type on your screen to do chores for you.
Paid Best for Coding Visit
Native multimodal intelligence with a 10M context window
Added Feb 5, 2026
Google's most powerful multimodal model, capable of processing hours of video, thousands of lines of code, or massive document sets in a single prompt
Why: Gemini 3 Ultra offers an unbeatable 10M token context window, allowing it to process entire project histories, hours of video, or massive codebases in a single prompt. Its native multimodal intelligence makes it the only model capable of 'seeing' and 'hearing' complex data sets with the same level of depth as it reads text, providing a unique advantage for large-scale data analysis.
Paid Best for Context Visit
The secure backbone for agentic AI applications
Added Feb 5, 2026
RANA 2
Why: The 'Security' play. As agents become autonomous, the RANA framework provides the essential safety and cost-optimization layer for enterprise deployment.
Enterprise Best for Security Visit
On-demand GPU cloud for serverless AI inference
Added Feb 5, 2026
RunPod provides globally distributed GPU instances and serverless endpoints for AI model inference and training
Why: The 'Scale' play. Its massive global GPU availability and sub-second cold starts make it the best choice for high-traffic AI applications.
Enterprise Best for Scaling Visit
The conversational search engine that replaced traditional search
Added Feb 5, 2026
Perplexity uses frontier LLMs to browse the web in real-time and provide cited, accurate answers to complex queries
Why: Perplexity AI is the 'Death of the Search Engine.' Instead of giving you a list of 10 links to click on, it just reads the whole internet for you and gives you a single, cited answer. It's like having a personal researcher who never sleeps.
Freemium Best for Research Visit
AI search engine for peer-reviewed scientific research
Added Feb 5, 2026
Consensus searches over 200 million scientific papers to provide evidence-based answers
Why: The 'Truth' layer for AI. It solves the hallucination problem in research by grounding every answer in peer-reviewed science.
Freemium Best for Science Visit
The agentic browser that takes action on the web
Added Feb 5, 2026
MultiOn is an AI agent that can use a web browser like a human
Why: The bridge to the 'Action' economy. It moves AI from 'talking' to 'doing' by interacting with the legacy web on behalf of the user.
Enterprise Best for Actions Visit
Instant high-quality 3D modeling from text and images
Added Feb 5, 2026
Tripo AI v3 generates high-fidelity 3D meshes with clean topology and PBR textures in seconds
Why: The fastest path to 3D. Its v3 engine produces meshes that are actually usable in production pipelines without massive manual cleanup.
Freemium Best for 3D Speed Visit
High-fidelity 3D asset generation from Luma Labs
Added Feb 5, 2026
Genie is Luma's specialized 3D generation engine
Why: The 'Midjourney' of 3D. It prioritizes aesthetic quality and texture detail, making it the best for visual-first 3D projects.
Freemium Best for 3D Detail Visit
The industry standard for cinematic AI video generation
Added Feb 6, 2026
Runway's Gen-4
Why: Runway Gen-4.5 is the 'Hollywood' of AI video. It creates the most realistic movies where characters move and look exactly like real people. It's the top choice for professional filmmakers because it gives them total control over the camera and the actors' expressions.
Paid Best for Filmmaking Visit
High-speed, high-realism video generation
Added Feb 5, 2026
Luma's Dream Machine v2 is a highly efficient video model known for its extreme realism and fast generation speeds
Why: Luma Dream Machine v2 is the 'Speed Demon' of AI video. It can turn a simple photo into a realistic 5-second video clip faster than almost any other tool. It's perfect for when you need to see your ideas come to life instantly.
Freemium Best for Realism Visit
The new gold standard for prompt adherence and text rendering
Added Feb 5, 2026
Black Forest Labs' FLUX
Why: FLUX.1 [pro] is the 'Master Artist' for AI images. Most AI tools are bad at writing words inside pictures, but Flux is perfect at it. It's the best tool for designers who need high-quality posters, logos, and photos that look 100% real.
Paid Best for Design Visit
The creative suite for physics-defying video effects
Added Feb 5, 2026
Pika 2
Why: Pika 2.0 is the 'Fun Lab' for AI video. It has special 'Pikaffects' that let you do crazy things like squish, melt, or explode objects in your videos. It's the best tool for making funny, viral videos for social media.
Freemium Best for Viral Content Visit
Production-ready 3D assets in under 60 seconds
Added Feb 5, 2026
Meshy v3 is the fastest text-to-3D and image-to-3D engine, producing high-topology meshes with PBR textures
Why: The bridge between AI and Game Engines. It generates usable, textured meshes that can be dropped directly into Unity or Unreal without manual cleanup.
Freemium Best for Game Dev Visit
The global infrastructure for open-source model deployment
Added Feb 5, 2026
Now integrated with Cloudflare's global network, Replicate allows you to run and fine-tune open-source models (Flux, Llama, Whisper) with a single API call and zero infrastructure management
Why: The 'GitHub' of model deployment. It democratizes access to the world's best open-source models with enterprise-grade scaling.
Enterprise Best for Open Source Visit
Generate and refine 3D assets from text or images
Added Feb 5, 2026
Generates 3D meshes from text prompts or images using AI-powered reconstruction
Why: Meshy AI provides the fastest professional speed-to-3D workflow, enabling artists to iterate from a simple text prompt or 2D image to a usable, textured mesh in under a minute. Its high-quality PBR texture generation and clean topology make it the most efficient tool for game developers and 3D prototypers looking to bypass manual modeling bottlenecks.
Freemium Best for 3D Assets Visit
Alibaba flagship video with joint audio and multilingual lip-sync
Added May 3, 2026
Alibaba's HappyHorse 1
Why: It addresses the hardest user complaint about AI video, convincing sound and lip-sync with motion, not only pixels. Strong fit when you need dialogue-forward clips or localized performances without a full audio post stack.
Paid Best for Audio+Video Visit
One multimodal model for text, vision, audio, and video reasoning
Added May 3, 2026
Nemotron 3 Nano Omni is NVIDIA's compact-but-capable multimodal stack for agentic workflows: one family of endpoints that accept text, images, audio, or video (depending on route) and return text answ...
Why: If your product roadmap says 'agents that see and hear the world,' Omni is built for that integration story, fewer moving parts than bolting Whisper + CLIP + LLM together by hand.
Paid Best for Agents Visit
Multi-image to production-grade 3D on next-gen Meshy
Added May 3, 2026
Meshy 6 continues Meshy's focus on fast, usable 3D assets with emphasis on multi-image conditioning: feed several views or references so the model better infers shape, materials, and proportions for g...
Why: Teams outgrew 'cool sculpt from one photo' and need consistent assets from multiple references, Meshy 6 is explicitly positioned for that workflow.
Freemium Best for Multi-Ref 3D Visit
Real-time virtual try-on in video
Added May 3, 2026
Lucy 2
Why: Most directories list generic video models; few spell out 'commerce motion' workflows, VTON fills that gap for teams selling apparel and accessories.
Paid Best for Fashion Video Visit
Fine-tuned control with adjustable inference
Added Feb 5, 2026
Generates images with adjustable inference steps and guidance scale using Flux 2 Flex model, featuring enhanced typography and text rendering capabilities
Why: Best control over generation parameters + superior text rendering, making it ideal for projects requiring precise control and accurate text in images.
Best for Control Visit
Alibaba's open-source MoE flagship with thinking modes
Added Apr 28, 2025
Qwen 3 is a 2025 open-weight Mixture-of-Experts model family from Alibaba Cloud, ranging from 0
Free Best for Open-Source Agents Visit
Alibaba's open coding-specialist model
Added Nov 12, 2024
Qwen 2
Free Best for Open Coding Visit
Alibaba's closed-API flagship before Qwen 3
Added Jan 28, 2025
Qwen 2
Paid Best for API Flagship Visit
Alibaba's strongest vision-language model
Added Jan 29, 2024
Qwen-VL-Max is a high-performance vision-language model from Alibaba, capable of understanding images, charts, and documents, and answering questions about them
Freemium Best for Vision-Language Visit
Open mathematical reasoning specialist
Added Aug 1, 2024
Qwen-Math is a family of open-weight models specialized for mathematical reasoning and problem solving, derived from Qwen 2
Free Best for Math Reasoning Visit
Earlier open-source Wan video generation model
Added Mar 1, 2025
Wan 2
Free Best for Open Video Generation Visit
Ultra-realistic FLUX.1 update with faster generation
Added Oct 2, 2024
FLUX
Paid Best for Premium Image Quality Visit
Advanced inpainting and outpainting FLUX model
Added Oct 2, 2024
FLUX
Paid Best for Image Editing Visit
Canny-edge-guided image generation and editing
Added Oct 2, 2024
FLUX
Paid Best for Structural Control Visit
Depth-map-guided image generation and editing
Added Oct 2, 2024
FLUX
Paid Best for Spatial Control Visit
Fast local FLUX.2 generation for personal hardware
Added Jun 1, 2025
FLUX
Free Best for Fast Local Generation Visit
Open-weight FLUX.2 for research and commercial use
Added Jun 1, 2025
FLUX
Free Best for Open Customization Visit
Open-source JSON-native text-to-image model built for controllable, enterprise-safe generation
Added Jun 26, 2025
BRIA FIBO is an 8B-parameter DiT text-to-image model trained on long structured JSON captions
Why: FIBO stands out for native JSON structured prompting and fully licensed training data, making it the strongest open-source choice for enterprises that need predictable, legally safe image generation.
Freemium Best for Controllable Image Generation Visit
Fast, lightweight FIBO pipeline designed for speed, efficiency, and on-prem deployment
Added Nov 11, 2025
BRIA FIBO Lite is a lightweight variant of the FIBO image generation pipeline
Why: FIBO Lite gives teams a FIBO-family option optimized for speed and data sovereignty, with a fully local deployment path that the full FIBO pipeline does not emphasize.
Freemium Best for Fast Local Image Generation Visit
High-accuracy background removal model trained on a licensed, professionally labeled dataset
Added Jul 24, 2025
BRIA RMBG 2
Why: RMBG 2.0 is a widely adopted, source-available background removal model with strong commercial licensing and a dedicated GitHub presence, filling a clear gap alongside BRIA's eraser tools.
Freemium Best for Background Removal Visit
Fast, low-cost Claude model with extended thinking and Computer Use
Added Oct 15, 2025
Anthropic's Haiku-tier model announced on October 15, 2025, and the first Haiku model to support both extended thinking and Computer Use
Why: Haiku 4.5 is the only current Haiku model and is the first to bring Computer Use and extended thinking to the low-cost tier.
Paid Best for Fast, Low-Cost Agents Visit
Balanced Sonnet model with major coding and agentic improvements
Added Sep 29, 2025
Anthropic's Sonnet-tier model announced on September 29, 2025, with major coding, instruction-following, and agentic improvements at the same price as Sonnet 4
Why: Sonnet 4.5 brought meaningful coding and agentic upgrades at the same price as Sonnet 4, making it a notable missing mid-tier entry.
Paid Best for Balanced Coding Agents Visit
First Claude model with the effort parameter and context compaction
Added Nov 24, 2025
Anthropic's Opus-tier model announced on November 24, 2025, introducing the effort parameter for balancing capability against cost, context compaction, and a deeper memory tool
Why: Opus 4.5 was the first Claude model to ship the effort parameter, an important capability evolution before Opus 4.6 and 4.7.
Enterprise Best for Cost-Capability Tradeoffs Visit
Frontier Opus model with higher-resolution vision and xhigh effort
Added Apr 16, 2026
Anthropic's frontier Opus-tier model announced on April 16, 2026, with substantial gains on the hardest coding tasks, higher-resolution vision input, a new xhigh effort level, file-system-based memory...
Why: Opus 4.7 introduced the xhigh effort level and file-system memory recall, making it a notable step between Opus 4.6 and Opus 4.8.
Enterprise Best for Hard Coding Tasks Visit
Limited-availability Mythos-class model without Fable 5 safety classifiers
Added Jun 9, 2026
Anthropic's Mythos-class model announced on June 9, 2026, shares the same capabilities as Claude Fable 5 without the safety classifiers
Why: Mythos 5 is a notable limited-availability variant of the Mythos-class tier, distinct from the generally available Fable 5.
Enterprise Best for Controlled Research Visit
High-volume DeepSeek inference with a 1M-token context window
Added Apr 24, 2026
DeepSeek V4-Flash is the efficient sibling of V4-Pro, offering a 1M-token context window and configurable thinking modes at a fraction of the API cost
Why: V4-Flash delivers the same 1M context and thinking modes as V4-Pro at roughly one-third the API cost, making it the practical default for most production workloads.
Freemium Best for High-Volume APIs Visit
The open-weight reasoning model that sparked the efficiency revolution
Added Jan 20, 2025
DeepSeek R1 is a 671B-parameter open-weight reasoning model that matches o1-class performance on math, code, and logic benchmarks through reinforcement learning on verifiable tasks
Why: R1 proved that open-weight models can match proprietary reasoning systems at a fraction of the cost, making it a landmark for reproducible AI research.
Freemium Best for Open Reasoning Visit
The 128K-context MoE flagship that introduced sparse attention
Added Dec 1, 2025
DeepSeek V3
Why: V3.2 introduced DeepSeek Sparse Attention and unified thinking modes, making it the architectural bridge that enabled the later 1M-context V4 family.
Freemium Best for Long-Context MoE Visit
Emotionally-aware multilingual text-to-speech across 29 languages
Added Aug 1, 2023
Produces natural, lifelike text-to-speech with rich emotional range and contextual understanding across 29 languages
Why: ElevenLabs' most emotionally-aware multilingual model, ideal for projects that need a consistent, expressive voice across many languages.
Freemium Best for Emotional Multilingual TTS Visit
Ultra-low-latency text-to-speech for real-time voice agents
Added Dec 1, 2024
Delivers high-quality speech synthesis with approximately 75ms latency across 32 languages, optimized for real-time voice agents, chatbots, interactive applications, and large-scale TTS processing
Why: The fastest ElevenLabs TTS model for production voice agents and real-time interactive experiences where latency matters.
Freemium Best for Real-Time Voice Visit
Real-time multilingual voice conversion that preserves emotion and content
Added Jun 1, 2024
Converts one voice into another while preserving intonation, emotion, accent, and spoken content across 29 languages
Why: A dedicated voice conversion model that keeps emotion, accent, and content intact across languages.
Freemium Best for Voice Conversion Visit
Generate custom synthetic voices from text descriptions
Added Jun 1, 2025
Creates entirely new synthetic voices from a text prompt describing the desired age, gender, accent, personality, and style, supporting 70+ languages
Why: Lets creators design unique voices from a written description, eliminating the need for recorded samples.
Freemium Best for Voice Design Visit
744B-parameter open-weight MoE flagship for agentic planning and execution
Added Feb 11, 2026
GLM-5 is Zhipu AI's (Z
Why: GLM-5 anchors the open-weight GLM line as a commercially-usable Chinese flagship with a permissive license and a strong reasoning profile.
Paid Best for Open-Weight Frontier Visit
MIT-licensed MoE flagship for 8-hour autonomous coding sessions
Added Apr 7, 2026
GLM-5
Why: GLM-5.1 is the open-weight coding release that made Z.ai competitive on long-horizon agentic work while remaining MIT-licensed for unrestricted commercial use.
Paid Best for Long-Horizon Coding Visit
Optimized GLM-5 variant for fast sequential task execution
Added Jun 1, 2026
GLM-5-Turbo is a tuned variant of the GLM-5 series that prioritizes lower latency and efficient sequential execution
Why: GLM-5-Turbo is the practical speed layer for the GLM-5 family, trading a small amount of peak capability for noticeably faster multi-step agent execution.
Paid Best for Fast Sequential Tasks Visit
Strong general-reasoning model with interleaved thinking
Added Dec 1, 2025
GLM-4
Why: GLM-4.7 is the cost-effective sweet spot for long-context reasoning and general-purpose agent work before stepping up to the GLM-5 series.
Paid Best for General Reasoning Visit
Mid-range coding and tool-calling model with 200K context
Added Sep 1, 2025
GLM-4
Why: GLM-4.6 gives developers a capable, lower-cost GLM option for coding and tool-calling agents without sacrificing the long context window.
Paid Best for Coding & Tool Calls Visit
Cost-efficient reasoning, coding, and agent model
Added Jul 1, 2025
GLM-4
Why: GLM-4.5-Air extends the GLM-4.5 family downward with a low-cost model that still handles coding, reasoning, and agent tasks.
Paid Best for Budget Reasoning Visit
Multimodal coding and visual-reasoning agent model
Added Jun 1, 2026
GLM-5V-Turbo is a vision-language variant of the GLM-5 family, built for multimodal coding, visual reasoning, and image-plus-text agent workflows
Why: GLM-5V-Turbo is the GLM family's main vision agent, letting coding and agent workflows reason over images and screenshots in the same long context.
Paid Best for Multimodal Coding Visit
Vision-language model for visual reasoning and UI replication
Added Dec 1, 2025
GLM-4
Why: GLM-4.6V is the practical vision tier for turning screenshots and images into working code or structured analysis.
Paid Best for Visual Reasoning Visit
Document parsing model for PDF and image OCR
Added Oct 1, 2025
GLM-OCR is a specialized GLM model for extracting structured Markdown from PDFs and images
Why: GLM-OCR fills a clear gap in the GLM family by turning scanned documents and PDFs into structured, usable text with layout awareness.
Paid Best for Document Parsing Visit
Free universal GLM model with a 200K context window
Added Apr 1, 2026
GLM-4
Why: GLM-4.7-Flash is the most capable free-tier GLM option, making long-context prototyping accessible without a subscription.
Free Best for Free General Use Visit
Free vision model for image understanding and document snapshots
Added Apr 1, 2026
GLM-4V-Flash is a free-tier vision model in the GLM family, offering zero-cost image understanding and document snapshot analysis with a 16K context window
Why: GLM-4V-Flash is the entry-level vision option for GLM, letting users test multimodal document understanding before upgrading to paid vision tiers.
Free Best for Free Vision Visit
Google's long-context multimodal flagship with up to 2M tokens
Added Feb 15, 2024
Gemini 1
Freemium Best for Long Context Visit
Fast, cost-efficient multimodal model with a 1M context window
Added May 21, 2024
Gemini 1
Freemium Best for Fast Multimodal Tasks Visit
Google's low-latency agentic model with native tool use
Added Dec 11, 2024
Gemini 2
Freemium Best for Agentic Apps Visit
Google's high-performance reasoning model with advanced coding
Added Mar 25, 2025
Gemini 2
Freemium Best for Complex Reasoning Visit
Google's photorealistic text-to-image model with text rendering
Added Dec 13, 2023
Imagen 2 is a diffusion-based text-to-image model developed by Google DeepMind
Paid Best for Realistic Images Visit
Google's high-quality 1080p video generation model
Added Dec 16, 2024
Veo 2 is a 2024 text-to-video and image-to-video model from Google DeepMind that produces 1080p cinematic clips with strong prompt adherence, camera control, and realistic motion
Paid Best for Cinematic Video Visit
Google's open multimodal model for research and developers
Added Mar 12, 2025
Gemma 3 is an open-weights family of multimodal models from Google, ranging from 1B to 27B parameters
Free Best for Open Multimodal Visit
xAI's long-context flagship with a 1M-token window
Added Apr 1, 2026
Grok 4
Why: Grok 4.3 is the sweet spot in xAI's lineup for anyone who needs a frontier model with a very large context window at a lower price than Grok 4.5.
Paid Best for Long-Context Work Visit
xAI's high-volume, 2M-context workhorse model
Added Nov 1, 2025
Grok 4
Why: Grok 4.1 Fast delivers one of the largest context windows in the family at the lowest price point, making it the default pick for bulk work.
Paid Best for Volume and Cost Efficiency Visit
xAI's 2M-context beta model with multi-agent capabilities
Added Feb 1, 2026
Grok 4
Why: Grok 4.20 remains notable as xAI's first multi-agent beta model with a 2M context window, even though newer 4.3/4.5 models now offer flagship alternatives.
Paid Best for Multi-Agent Beta Work Visit
xAI's fast, cheap coding specialist model
Added Aug 1, 2025
Grok Code Fast 1 is a lightweight coding-focused model from xAI with a 256K context window
Why: Grok Code Fast 1 is the practical, low-cost coding specialist in xAI's family for developers who want Grok reasoning without flagship pricing.
Paid Best for Fast Coding Assistance Visit
Tencent's 389B-parameter open-source MoE language model
Added Nov 4, 2024
Open-source Transformer-based Mixture-of-Experts language model with 389 billion total parameters and 52 billion active parameters
Why: Largest open-source Transformer-based MoE model from Tencent, ideal for researchers and builders who want to self-host a capable long-context LLM.
Free Best for Open-Source LLM Workloads Visit
Tencent's Mamba-powered deep-thinking reasoning model
Added Mar 21, 2025
Hybrid Mamba-Transformer MoE reasoning model released March 2025, built on Hunyuan TurboS with 52 billion active parameters and a 256K context window
Why: One of the first ultra-large Mamba-Transformer MoE reasoning models, offering strong benchmark scores and a 256K context window.
Freemium Best for Reasoning Visit
Tencent's fast, cost-efficient flagship Hunyuan model
Added Jan 10, 2026
A 200K-context open-weight Hunyuan model optimized for speed while maintaining strong performance on general chat, coding, and agentic tasks
Why: Speed-optimized Hunyuan flagship with a 200K context window and strong price/performance for production APIs.
Freemium Best for Speed Visit
Tencent's general-purpose instruction-tuned Hunyuan 2.0 model
Added Jun 20, 2025
Open-weight instruction-tuned variant of Tencent's Hunyuan 2
Why: Versatile instruction-tuned Hunyuan model balancing capability and context for a wide range of tasks.
Freemium Best for General-Purpose Chat Visit
The deep-thinking variant of Hunyuan 2.0
Added Sep 15, 2025
Open-weight reasoning variant of Hunyuan 2
Why: Hunyuan 2.0's reasoning mode for tasks that benefit from longer thought chains.
Freemium Best for Reasoning Visit
Tencent's efficient small-scale MoE instruct model
Added Nov 15, 2025
A compact open-weight MoE instruction model with a 131K context window, listed as a cost-efficient everyday Hunyuan option via OpenRouter and Tencent Cloud
Why: Smallest listed Hunyuan instruct model, making it attractive for budget-conscious long-context deployments.
Freemium Best for Cost-Efficient Inference Visit
Tencent's latest open-source MoE flagship with tool use
Added Apr 22, 2026
Open-weight preview of Hunyuan 3 (Hy3), a 295B-parameter MoE model with 21B active parameters and a 262K context window
Why: Tencent's strongest open-source Hunyuan model to date, with competitive coding and agentic benchmarks.
Freemium Best for Coding and Agents Visit
Tencent's latest open-source high-fidelity 3D asset generator
Added Jun 13, 2025
Open-source system for generating high-resolution textured 3D assets from text or images
Why: Current open-source Hunyuan 3D pipeline with professional texture, PBR, and Blender integration.
Free Best for Production 3D Assets Visit
Realistic images, flexible styles, and reliable typography in one prompt
Added Aug 1, 2024
Generates photorealistic and stylized images from text prompts with a major leap in realism, prompt adherence, and text rendering over the first Ideogram model
Why: Ideogram 2.0 was the release that made Ideogram a serious alternative to Midjourney for realistic, text-heavy marketing imagery before V3 arrived.
Freemium Best for Realistic Marketing Images Visit
Fast, low-cost generation for rapid creative exploration
Added Sep 1, 2024
A speed-optimized variant of Ideogram 2
Why: Ideogram 2a gives creators a faster, cheaper way to produce the same text-in-image style when iteration speed matters more than pixel-perfect quality.
Freemium Best for Fast Iteration Visit
Moonshot's open-weight multimodal generalist with agent swarms
Added Jan 27, 2026
Kimi K2
Why: Kimi K2.5 was Moonshot's first widely available open-weight multimodal generalist and remains a notable reference point for the K2 family before K2.6 and K3 arrived.
Freemium Best for Open Multimodal Agents Visit
Moonshot's open-weight multimodal successor with long-context coding stability
Added Apr 21, 2026
Kimi K2
Why: Kimi K2.6 improves on K2.5 with stronger long-context coding and is a practical open-weight alternative for teams that want multimodal agents without the cost of closed frontier models.
Freemium Best for Long-Context Coding Visit
Faster inference variant of Kimi's coding specialist
Added Jun 12, 2026
Kimi K2
Why: Kimi K2.7 Code Highspeed is the latency-optimized version of an already strong coding model, making it a good pick for interactive coding agents and live pair-programming workflows.
Freemium Best for Fast Coding Visit
Kling's first widely available video generation model
Added Jun 6, 2024
Kling 1
Freemium Best for Early Kling Video Visit
Kling's standard model for cinematic video
Added Oct 1, 2024
Kling 2
Freemium Best for Cinematic Standard Visit
Improved physics and expressive movement in Kling video
Added Jan 1, 2025
Kling 2
Freemium Best for Expressive Motion Visit
Premium tier of Kling 3.0 with best quality
Added Apr 1, 2025
Kling 3
Paid Best for Premium Quality Visit
Kling's image generation model with style control
Added Mar 1, 2025
Kling Image 2
Freemium Best for Styled Images Visit
Controllable cinematic video model with multi-keyframe direction and motion transfer
Added Jul 15, 2026
Luma Ray 3
Why: Ray 3.2 gives professional teams frame-level control over video generation, including motion transfer and EXR export, making it a strong contender for production pipelines.
Freemium Best for Cinematic Control Visit
Multimodal reasoning model that generates brand-consistent images and edits
Added Jun 15, 2026
Luma Uni-1
Why: Uni-1.1 ties a reasoning model directly to pixel generation, making it unusually good at following brand references and complex creative direction in images.
Freemium Best for Brand-Consistent Images Visit
Fast text- and image-to-3D for concept exploration
Added Jan 1, 2024
Meshy 5 is a 2024-generation model that turns text prompts or reference images into textured 3D meshes in about 45 seconds
Why: It is the fast-iteration sibling in Meshy's current model family, still available for creators who want usable concepts in under a minute.
Freemium Best for Fast Iteration Visit
Clean, controllable game-ready topology in ~10 seconds
Added Jul 21, 2026
Smart Topology is Meshy's 2026 in-house model that generates 3D models with native, cleanly structured geometry and a controllable polygon count from 100 to 15,000
Why: It brings explicit polygon budgets and clean topology to AI-generated meshes, making it Meshy's most game-engine-ready model.
Freemium Best for Game-Ready Topology Visit
Meta's open-weight flagship with native multimodal reasoning
Added Apr 5, 2025
Llama 4 Maverick is Meta's flagship open-weight model, released in April 2025 as part of the Llama 4 family
Why: Maverick is the top open-weight model Meta actually ships today, with strong multimodal reasoning and a practical API ecosystem, making it the default choice for open Llama deployments.
Free Best for Open Multimodal Reasoning Visit
Long-context, efficient open multimodal model for edge and single-GPU use
Added Apr 5, 2025
Llama 4 Scout is Meta's efficient Llama 4 variant, released in April 2025
Why: Scout is notable for its extreme 10M-token context window and efficient single-GPU deployment, making it the standout open model for very long documents and memory-heavy applications.
Free Best for Long Context Visit
Efficient 70B open model matching 405B quality
Added Dec 6, 2024
Llama 3
Why: Llama 3.3 is the practical sweet spot in the Llama family: it gives users near-frontier open-model quality in a 70B package that is far cheaper to host and fine-tune than the 405B model.
Free Best for Efficient Open LLMs Visit
The first frontier-scale open-weight language model
Added Jul 23, 2024
Llama 3
Why: Llama 3.1 405B remains a landmark open release: it proved open weights could compete with proprietary frontier models and still serves as a high-quality baseline for research and synthetic-data generation.
Free Best for Frontier Open Research Visit
Low-code platform for building and managing custom AI agents
Added Nov 1, 2023
Microsoft Copilot Studio is a graphical, low-code SaaS platform for designing custom AI agents and agentic workflows
Why: The tool that turns the Microsoft Copilot ecosystem into a custom agent platform without requiring heavy coding.
Enterprise Best for Custom Agents Visit
Recursive self-improvement language model for real-world engineering
Added May 1, 2026
MiniMax M2
Why: MiniMax M2.7 is the current production language model below M3 and is explicitly listed as beginning recursive self-improvement, making it a notable addition to the family.
Freemium Best for Engineering Tasks Visit
Same M2.7 performance with significantly faster inference
Added May 1, 2026
MiniMax M2
Why: The Highspeed variant is a current, actively promoted option for developers who need M2.7 capability with lower latency.
Freemium Best for Low-Latency Coding Visit
Ultra-realistic multilingual text-to-speech with sound tags
Added Jun 1, 2026
MiniMax Speech 2
Why: Speech 2.8 HD is the current quality-tier MiniMax voice model, replacing the earlier Speech 2.6 / Speech-02 series.
Freemium Best for Realistic Speech Visit
Fast multilingual text-to-speech with natural flow
Added Jun 1, 2026
MiniMax Speech 2
Why: Speech 2.8 Turbo is the current speed-tier MiniMax voice model, distinct from the HD quality variant.
Freemium Best for Real-Time TTS Visit
Music generation with humanized vocals and elevated sound
Added Jun 1, 2026
MiniMax Music 3
Why: Music 3.0 is the current MiniMax music generation model, replacing the legacy Music 2.0 entry already in the directory.
Freemium Best for Vocal Music Visit
Mistral's flagship open-weight multimodal frontier model
Added Apr 15, 2026
A 675B-parameter sparse mixture-of-experts model with 41B active parameters and a 262K context window, released under Apache 2
Why: Mistral Large 3 is one of the most capable permissive open-weight models available, offering frontier performance with the deployment flexibility of Apache 2.0 licensing.
Freemium Best for Open-Weight Frontier Visit
Mistral's mid-tier workhorse for reasoning, coding, and instruction
Added May 22, 2026
A mid-tier model that balances performance and cost, optimized for instruction following, reasoning, and coding
Why: Mistral Medium 3.5 delivers strong performance at a lower cost than the flagship, making it the sensible default for most business and development workloads.
Freemium Best for Everyday Workloads Visit
Unified open-source small model for chat, reasoning, vision, and coding
Added May 1, 2026
A 119B-parameter MoE model with 6B active parameters and a 256K context window, released under Apache 2
Why: Small 4 packs flagship-class reasoning, vision, and coding into a single open-source model that is efficient enough for high-throughput and local deployments.
Freemium Best for Efficient Open Multimodal Visit
Mistral's code-specialist model with fill-in-the-middle support
Added Aug 1, 2025
A code generation model optimized for latency-sensitive fill-in-the-middle completion and chat, supporting 80+ programming languages
Why: Codestral 25.08 improves accepted completions and reduces runaway generations, making it a strong open-weight option for production IDE assistants.
Freemium Best for IDE Code Completion Visit
Mistral's edge family of small, dense open-source models
Added Apr 15, 2026
A family of 3B, 8B, and 14B parameter dense models released under Apache 2
Why: Ministral 3 brings Mistral's open-weight lineage to edge devices, offering a strong 14B reasoning option and smaller variants for local and on-device use.
Freemium Best for Edge Deployment Visit
Mistral's open-weight speech understanding and TTS models
Added Jul 1, 2025
A family of open-weight speech models including a 24B production variant and a 3B edge variant, released under Apache 2
Why: Voxtral offers open-weight speech understanding and synthesis at a fraction of the cost of proprietary alternatives, making it practical for production voice agents.
Freemium Best for Voice AI Visit
Compact 30B open-weight model with configurable reasoning for agents
Added Jun 4, 2026
A 30B total / 3B active parameter hybrid Mamba-2 + Transformer MoE language model built for efficient on-device and edge agentic tasks
Why: The smallest open-weight member of the Nemotron 3 family, giving teams frontier-style reasoning and tool-use without data-center hardware.
Free Best for Efficient Agents Visit
120B open-weight hybrid MoE for efficient multi-agent reasoning
Added Jun 4, 2026
A 120B total / 12B active parameter hybrid Mamba-Transformer MoE language model with LatentMoE, multi-token prediction, and native NVFP4 pretraining
Why: Fills the gap between Nano and Ultra with a strong efficiency-to-accuracy ratio for agentic orchestration and latency-sensitive serving.
Free Best for Multi-Agent Efficiency Visit
Layout-aware document parsing that goes beyond OCR
Added Jun 1, 2025
A general-purpose document parsing model that overcomes traditional OCR limitations by understanding complex page layouts
Why: Turns messy documents into clean structured data, improving downstream RAG and agent pipelines with layout-aware extraction.
Free Best for Document Intelligence Visit
Multimodal 4B safety model for text and image moderation
Added Jun 4, 2026
A 4B-parameter multimodal, multilingual small language model designed as a robust content-safety moderator
Why: A compact, open safety model that can enforce both standard and custom content policies with reasoning traces for safer deployments.
Free Best for Content Safety Visit
NVIDIA-aligned 253B Llama 3.1 for helpfulness and instruction following
Added Dec 1, 2024
A 253B-parameter variant of Llama 3
Why: NVIDIA's largest aligned Llama collaboration, offering a strong open-weight alternative for teams already standardizing on Llama architectures.
Free Best for Aligned Llama Performance Visit
NVIDIA-aligned 49B Llama 3.1 for balanced performance
Added Dec 1, 2024
A 49B-parameter variant of Llama 3
Why: A mid-size aligned Llama model that balances capability and deployment cost for teams using NVIDIA tooling.
Free Best for Balanced Llama Deployment Visit
NVIDIA-aligned 8B Llama 3.1 for efficient inference
Added Dec 1, 2024
An 8B-parameter variant of Llama 3
Why: A compact, NVIDIA-aligned Llama model for teams that need HelpSteer-tuned instruction following on limited hardware.
Free Best for Efficient Aligned Llama Visit
Lightweight, real-time search model for cited answers at low cost
Added Jan 23, 2024
Perplexity's base Sonar model pairs live web search with a compact LLM to deliver fast, citation-backed answers
Why: Sonar is the affordable, fast entry point to Perplexity's live-search API and is widely used as the default model for citation-backed Q&A.
Paid Best for Everyday Search Visit
Advanced search model with deeper reasoning and richer citations
Added Nov 1, 2024
Sonar Pro uses a more capable model and expanded search context to answer complex questions with detailed, source-backed responses
Why: Sonar Pro adds the depth and reliability needed for serious research while remaining accessible through Perplexity's API and Pro app tier.
Paid Best for Deep Research Visit
Chain-of-thought reasoning model for multi-step logical analysis
Added Jan 21, 2025
Sonar Reasoning Pro exposes explicit chain-of-thought reasoning to solve complex, multi-step problems with transparent intermediate steps
Why: Sonar Reasoning Pro is Perplexity's option for users who need transparent, step-by-step reasoning rather than just a final answer.
Paid Best for Complex Reasoning Visit
Autonomous research agent that performs multi-source deep dives
Added Feb 14, 2025
Sonar Deep Research conducts autonomous, multi-step research across many web sources and synthesizes a comprehensive, cited report
Why: Sonar Deep Research is the most agentic member of the Sonar family, capable of running lengthy, autonomous research workflows and producing publishable reports.
Enterprise Best for Autonomous Research Visit
Pika's first public video generation model
Added Nov 28, 2023
Pika 1
Freemium Best for First Pika Video Visit
Pika's upgrade with improved motion and effects
Added Oct 1, 2024
Pika 1
Freemium Best for Pikaffects Visit
Pika's refined model with stronger realism and camera control
Added Jun 1, 2025
Pika 2
Freemium Best for Realistic Pika Video Visit
Second-generation designer-first image generation model
Added Mar 13, 2024
Recraft V2 is the second-generation image generation model released by Recraft in March 2024
Why: Recraft V2 was the first generational upgrade that explicitly positioned Recraft as a designer-first model with strong style and anatomy control.
Freemium Best for Design Assets Visit
Recraft's most advanced image model with photorealistic, vector, and utility variants
Added May 14, 2026
Recraft V4
Why: Recraft V4.1 is the current flagship model, offering more natural photorealism, refined illustration quality, and dedicated Utility and Vector variants for production design workflows.
Freemium Best for Photorealistic Design Visit
Runway's first generation of text- and image-to-video
Added Mar 1, 2023
Runway Gen-2 is an earlier-generation video foundation model that generates short video clips from text prompts or images
Freemium Best for Early AI Video Visit
Faster, cheaper Gen-3 Alpha for rapid video iteration
Added Jul 31, 2024
Runway Gen-3 Alpha Turbo is a faster and more cost-efficient variant of Gen-3 Alpha, designed for creators who need to iterate quickly on video concepts without sacrificing too much quality
Freemium Best for Fast Iteration Visit
Runway's next-generation model for consistent characters and camera
Added Apr 1, 2025
Runway Gen-4 is a 2025 video generation model that emphasizes consistent characters, objects, and environments across multiple clips, along with advanced camera control and world consistency
Paid Best for Consistent Worlds Visit
Image generation model with strong style control
Added Oct 1, 2024
Runway Frames is a dedicated image generation model from Runway, designed to create stylized images with strong consistency and to serve as the starting frame for video generations
Freemium Best for Style-Locked Images Visit
Performance-driven character animation from video
Added Nov 1, 2024
Runway Act-One is a tool that transfers an actor's facial performance and expressions onto a generated character using video input, enabling expressive character animation without motion-capture hardw...
Paid Best for Performance Transfer Visit
Fast, inference-efficient 1080p video with native multi-shot storytelling
Added Jun 11, 2025
Seedance 1
Why: Seedance 1.0 established the foundation for ByteDance's video generation line with a strong emphasis on inference speed and native multi-shot coherence.
Freemium Best for Fast 1080p Video Visit
Cinematic audio-video joint generation with lip-sync and dialect support
Added Dec 16, 2025
Seedance 1
Why: Seedance 1.5 pro moved the family from silent video to native audio-visual generation, with strong lip-sync and dialect support that makes it practical for short-form drama and advertising.
Freemium Best for Audio-Visual Sync Visit
The foundational open-source text-to-image model
Added Oct 20, 2022
Stable Diffusion 1
Free Best for Foundation Ecosystem Visit
High-resolution open-source image generation
Added Jul 26, 2023
Stable Diffusion XL (SDXL) is a 2023 open-source text-to-image model that generates higher-quality, higher-resolution images than SD 1
Free Best for High-Resolution Open Images Visit
Fast one-step SDXL for real-time generation
Added Nov 28, 2023
Stable Diffusion XL Turbo is a distilled, fast variant of SDXL that can generate images in a single step or a few steps
Free Best for Fast Open Images Visit
Stability AI's first multimodal-diffusion Transformer image model
Added Jun 12, 2024
Stable Diffusion 3 is a 2024 text-to-image model from Stability AI based on a Multimodal Diffusion Transformer architecture
Free Best for Text-in-Image Visit
Efficient SD3 variant for consumer hardware
Added Jun 12, 2024
Stable Diffusion 3 Medium is a 2B-parameter version of SD3 designed to run well on consumer GPUs
Free Best for Local SD3 Visit
Stability AI's largest 3.5 model with best quality
Added Oct 22, 2024
Stable Diffusion 3
Free Best for SD3.5 Quality Visit
Earlier generation of Tripo's text- and image-to-3D pipeline
Added Jun 1, 2024
Tripo 2
Freemium Best for Generation Pipeline Visit
Mid-generation upgrade between Tripo 3 and 4
Added Sep 1, 2025
Tripo 3
Freemium Best for Speed-Quality Balance Visit
Tripo's latest high-fidelity 3D generation model
Added Dec 1, 2025
Tripo 4
Paid Best for Fidelity Visit
Anonymous 1M-context reasoning model available free through OpenRouter
New this month Added Aug 20, 2026
Ox Alpha is a stealth AI model that appeared on OpenRouter and OpenCode on 20 August 2026
Why: The combination of a one-million-token context window, multimodal inputs, free pricing during the preview, and rapid adoption by coding-agent builders makes it worth tracking even before its creator is known. Within a day of launch, coding agents had pushed billions of tokens through it, suggesting real production interest rather than curiosity traffic. Preliminary independent DeepSWE testing also places it ahead of Claude Fable 5 and GPT-5.6 Sol on a small task subset.
Free Best for Anonymous Preview Visit
Design-forward image generation (logos, vectors, assets)
Added Feb 5, 2026
Generates design assets including logos, vectors, and brand visuals with clean, usable outputs
Why: Great for design assets when you want clean, usable outputs with vector-style graphics and brand-ready visuals.
Freemium Best for Design Visit
Context-aware image generation and editing
Added Feb 5, 2026
Generates and edits images with context awareness for better coherence using Flux Kontext model
Why: Context-aware generation for more coherent results, making it superior for image editing and variation tasks requiring consistency.
Best for Editing Visit
Google's coding workhorse, three weeks after 3.6 Flash
New this month Added Aug 13, 2026
Gemini 3
Why: It is the cheapest route to a current-generation Google coding model. Introductory pricing runs at half the standard Flash rate until the end of 2026, and the published jump over 3.6 Flash is large enough to matter on exactly the agentic and web-development work Flash-tier models are usually bought for.
Freemium Best for Coding Value Visit
Open-source image generation with flexibility
Added Feb 5, 2026
Generates images from text with open-source flexibility and community support using Stable Diffusion 3
Why: Open-source standard with extensive customization options, making it the foundation for many custom image generation workflows.
Free Best for Open Source Visit
Latest Wan for image variations and editing
Added Feb 5, 2026
Generates image variations and edits using Wan 2
Why: Latest Wan iteration for I2I with improved quality, representing the current state-of-the-art in Wan's image-to-image capabilities.
Best for Variations Visit
High-fidelity object removal from images
Added Feb 5, 2026
Removes unwanted objects from images with high fidelity and minimal artifacts using BRIA's advanced inpainting technology
Why: Best-in-class object removal with clean results, making it the top choice for professional image cleanup and editing workflows.
Best for Editing Visit
Microsoft's unified AI model family from Build 2026
Added Jul 7, 2026
Microsoft announced a family of MAI-branded models at Build 2026, including MAI-Thinking-1 for reasoning, MAI-Image-2
Why: The MAI family gives Microsoft a cohesive, enterprise-ready AI stack. For organizations already using Microsoft services, these models reduce friction by running inside familiar tools rather than requiring separate platforms.
Enterprise Best for Microsoft Ecosystem Visit
Microsoft's advanced 3D generation from text or images
Added Feb 5, 2026
Microsoft TRELLIS generates high-quality 3D models from text prompts or reference images using a unified Structured LATent (SLAT) representation
Why: Microsoft's state-of-the-art 3D generation model with best-in-class quality for both text-to-3D and image-to-3D workflows. Open-source availability and NVIDIA integration make it ideal for professional 3D asset creation.
Free Best for 3D Assets Visit
Object removal from video with high fidelity
Added Feb 5, 2026
Removes unwanted objects from video frames with high fidelity and temporal consistency using BRIA's video inpainting technology
Why: Best video object removal with frame-to-frame consistency, providing the most reliable video cleanup capabilities available.
Best for Editing Visit
Hosted Alibaba TTS across 16 languages, in a fast tier and a fidelity tier
Added Jul 21, 2026
Qwen-Audio-3
Why: The two-tier split is the useful part: most TTS vendors make you pick latency or fidelity across the whole account, and this exposes both behind one API. The 300ms-level first-packet latency on the Flash tier is fast enough for live conversational use.
Paid Best for Multilingual Speech Visit
Relight and recamera videos
Added Feb 5, 2026
Allows users to relight and recamera their videos with AI-powered adjustments using LightX Recamera technology
Why: Unique relighting + camera control for video post-production, offering capabilities not available in standard video editing tools.
Best for Editing Visit
Advanced video editing and effects
Added Feb 5, 2026
Provides video editing, effects, and generation capabilities with advanced control using Runway's Gen-3 Alpha model
Why: Runway's latest generation model with enhanced editing features, representing the cutting edge of integrated video generation and editing.
Freemium Best for Editing Visit
Advanced AI music generation with high-quality compositions
Added Feb 5, 2026
Generates complete musical compositions from text prompts using advanced AI techniques
Why: Top-tier music generation model with advanced composition capabilities, producing professional-quality music suitable for commercial use.
Best for Music Visit
High-quality music and sound effects generation
Added Feb 5, 2026
Generates high-quality music and sound effects from text prompts using StabilityAI's latest audio model
Why: StabilityAI's flagship audio model combining music and sound effects generation in one powerful tool, ideal for comprehensive audio production workflows.
Best for Music Visit
Multilingual text-to-speech with natural voice synthesis
Added Feb 5, 2026
Converts text to natural-sounding speech with multilingual support across numerous languages and voices
Why: Industry-leading TTS with exceptional voice quality and multilingual capabilities, making it the go-to choice for professional voice synthesis.
Best for Voice Visit
Open image generation ecosystem (model + tools)
Added Feb 5, 2026
Generates and edits images via an open model ecosystem including Stable Diffusion models and community tools
Why: Core ecosystem for customizable image workflows with open-source flexibility and extensive community support.
Best for Control Visit
Google's latest music generation model
Added Feb 5, 2026
Generates high-quality music from text prompts using Google's latest Lyria 2 model
Why: Google's cutting-edge music model representing the latest advances in AI music generation, with superior quality and versatility.
Best for Music Visit
CD-quality music with superior vocals
Added Feb 5, 2026
Generates CD-quality music from lyrics and style descriptions with superior vocal clarity and creative instrumentation
Why: Highest quality music generation with exceptional vocal production, making it ideal for commercial music creation requiring professional audio standards.
Best for Music Visit
Advanced sound effects generation
Added Feb 5, 2026
Generates professional-grade sound effects from text descriptions using ElevenLabs' advanced sound effects model
Why: ElevenLabs' latest sound effects model with superior quality and realism, ideal for professional audio production requiring high-fidelity SFX.
Best for SFX Visit
Fast Flux variant for rapid image generation
Added Feb 5, 2026
Generates high-quality images from text prompts using Black Forest Labs' Flux 1 schnell (fast) variant
Why: Fastest Flux variant maintaining top-tier quality, perfect for workflows requiring speed without compromising on image fidelity.
Best for Speed Visit
Google's high-quality text-to-image model
Added Feb 5, 2026
Generates realistic, high-quality images from text prompts using Google's Imagen 3 model
Why: Google's flagship image generation model with state-of-the-art quality and photorealism, representing one of the best text-to-image systems available.
Best for Quality Visit
Vector art and brand-style image generation
Added Feb 5, 2026
Generates long texts, vector art, and images in brand style using Recraft V3
Why: SOTA model excelling at vector art and brand consistency, making it unique for design workflows requiring precise style control and typography.
Best for Design Visit
Exceptional typography and text rendering
Added Feb 5, 2026
Generates high-quality images, posters, and logos with exceptional typography handling and realistic outputs
Why: Best-in-class typography rendering makes it the top choice for designs requiring text integration, logos, and marketing materials with readable text.
Best for Typography Visit
Development Flux for advanced control
Added Feb 5, 2026
Generates high-quality images from text prompts using Black Forest Labs' Flux 1 development version
Why: Development version offering advanced control and experimental features, ideal for developers and power users requiring maximum customization.
Best for Developers Visit
Quick text rendering for marketing graphics
Added Feb 5, 2026
Generates images optimized for quick, high-quality text rendering, making it suitable for creating marketing graphics with typography, UI mockups, and social media posts with captions
Why: Specialized for marketing graphics and text-heavy designs, making it the ideal choice for social media and UI mockup generation requiring readable text.
Best for Marketing Visit
Multilingual text rendering and photorealism
Added Feb 5, 2026
Generates images with multilingual text rendering and photorealism using a 6B parameter model optimized for deployment efficiency
Why: Unique multilingual text rendering capabilities make it essential for global marketing and content creation requiring text in multiple languages.
Best for Multilingual Visit
7B multimodal model for text and images
Added Feb 5, 2026
A 7B parameter multimodal model developed by ByteDance-Seed, capable of generating both text and images
Why: Unique multimodal capabilities combining text and image generation with editing, making it versatile for complex content creation workflows requiring multiple modalities.
Best for Multimodal Visit
Photorealistic Flux with LoRA fine-tuning
Added Feb 5, 2026
Generates photorealistic images from text prompts using Black Forest Labs' Flux model enhanced with Realism LoRA (Low-Rank Adaptation)
Why: Unique photorealistic variant of Flux with LoRA fine-tuning, offering specialized realism capabilities that complement the base Flux models for professional photography-style generation.
Best for Realism Visit
Customizable Flux with LoRA fine-tuning
Added Feb 5, 2026
Generates images from text prompts using Black Forest Labs' Flux model with LoRA (Low-Rank Adaptation) support for custom style fine-tuning
Why: LoRA-enabled Flux variant offering customizable style fine-tuning, making it ideal for specialized use cases requiring consistent character generation or specific artistic styles.
Best for Customization Visit
Multilingual text-to-speech with streaming
Added Feb 5, 2026
Converts text to natural-sounding speech using MiniMax's advanced TTS technology
Why: Comprehensive multilingual TTS solution with extensive voice library and streaming support, making it ideal for applications requiring real-time, multilingual voice synthesis across diverse use cases.
Best for Multilingual Visit
Open-source MoE LLM with strong Chinese NLP and multimodal capabilities
Added Jan 1, 2026
Baidu ERNIE 4
Why: Leading Chinese LLM with strong multilingual capabilities, open-source availability, and cost-efficient MoE architecture.
Freemium Best for Chinese Visit
Advanced multilingual LLM with enhanced reasoning and long-context support
Added Jan 1, 2026
GLM-4
Why: Advanced Chinese LLM with strong multilingual capabilities, efficient inference, and comprehensive deployment options.
Freemium Best for Multilingual Visit
Z.ai's post-trained coding and agentic model on the GLM-5.2 base
New this month Added Aug 14, 2026
GLM-5
Why: The Terminal-Bench 3.0 score moved from 4.6% to 28.3%, DeepSWE v1.1 from 46.2% to 66.9%, and CyberGym from 77.2% to 84.5% — all on the same base as GLM-5.2. That is a real signal about post-training returns, even if most numbers are vendor-run and weights are not yet released. It is also reported as one of the fastest models in its class, at roughly 115 tokens per second.
Paid Best for Post-Training Gains Visit
Open-source text-to-3D motion model with 200+ motion categories and production-ready exports
Added Jan 1, 2026
Hymotion 1
Why: Tencent's cutting-edge open-source text-to-3D motion model with production-ready output and extensive motion category support.
Free Best for 3D Motion Visit
Meta's closed-weight agentic model, and its first paid model API
New this month Added Aug 4, 2026
Muse Spark 1
Why: This is the release where Meta stopped giving models away. Muse Spark is closed, metered and sold through Meta's own API, and it lands at rank 15 on the independent index while undercutting comparable models on price. Worth tracking for that reason alone if your stack assumed Meta meant open weights.
Freemium Best Value for Agentic Multimodal Work Visit
NVIDIA's 550B open-weights reasoning model, built for inference speed
New this month Added Aug 4, 2026
Nemotron 3 Ultra is the largest member of NVIDIA's Nemotron 3 family, released 4 June 2026 at Computex
Why: The architecture is optimised for throughput rather than peak benchmark score, and it shows: over 400 output tokens per second at 550B parameters. It is also the most openly documented release at this scale, because publishing the datasets and post-training recipes lets you actually reproduce and extend the model instead of just running it.
Free Best for Open-Weight Throughput Visit
Autonomous AI agent for complex multi-step workflows and research automation
Added Jan 1, 2026
Manus is an autonomous AI agent developed by Butterfly Effect Pte
Why: Pioneering autonomous AI agent platform with proven real-world task execution capabilities, now backed by Meta's resources.
Enterprise Best for Automation Visit
Multimodal model generating image, video and audio from one set of weights
New this month Added Aug 4, 2026
FLUX 3 is Black Forest Labs' multimodal foundation model, announced 23 July 2026
Why: The first credible attempt to collapse image, video and audio generation into a single model rather than a pipeline of separate ones, from the team behind the most widely self-hosted open image models. Access is the catch: video is gated early-access and the open-weight release has not shipped, so treat availability as limited until FLUX 3 Dev lands.
Paid Best Multimodal Generation Visit
753B open-weight MoE coding model with a 1M-token context, MIT licensed
New this month Added Aug 4, 2026
GLM-5
Why: The strongest open-weight coding model published to date: 62.1 on SWE-bench Pro against GPT-5.5's 58.6, and 81.0 on Terminal-Bench 2.1, at roughly a sixth of GPT-5.5's API price. The MIT licence carries no regional restrictions, so the weights can genuinely be self-hosted commercially, which is the reason to choose it over a closed model of similar strength.
Freemium Best Open-Weight Coder Visit
Thinking Machines' 975B Apache-2.0 model that takes text, images and audio natively
New this month Added Aug 4, 2026
Inkling is the first open-weights model from Thinking Machines Lab, released 15 July 2026 under Apache 2
Why: It is the first roughly trillion-parameter open-weights model that takes audio and images natively rather than through a bolted-on encoder, and Apache 2.0 means the weights can be used commercially without asking anyone. Thinking Machines is candid that this is not the strongest model available but a base worth customising, which is a more honest pitch than most open releases make.
Free Best Open-Weight Multimodal Base Visit