BEST FOR • CURATED

Best AI Tools for AI Multimodal Reasoning

Best for AI Multimodal Reasoning

We've curated 46 top AI tools specifically selected for ai multimodal reasoning use cases. Each tool is evaluated for quality, reliability, and unique capabilities that make it well-suited for ai multimodal reasoning workflows.

WHY THESE TOOLS

These tools are selected because they excel at ai multimodal reasoning. When choosing, consider:

  • How the tool's specific features align with your ai multimodal reasoning needs
  • Whether the tool offers the right balance of quality, speed, and cost for your use case
  • Integration capabilities if you need to incorporate into existing workflows
  • Scalability for your production requirements
RESULTS
46 tools • curated
Anthropic's Mythos-class creative model
Added Jul 7, 2026
Claude Fable 5 is Anthropic's Mythos-class model released on June 9, 2026, focused on creative writing, worldbuilding, and narrative depth. It was suspended from distribution on June 12, 2026, under US export controls, making it a limited-availability release.
Why: Claude Fable 5 is notable as Anthropic's most experimental creative model. Even with limited availability, it represents an interesting direction for AI-assisted fiction and long-form creative work.
Paid Best for Creative Writing Visit
Anthropic's frontier model, currently first on the Artificial Analysis Intelligence Index
Added Aug 4, 2026
Claude Opus 5 is Anthropic's flagship model, released 24 July 2026 with a 1M-token context window and five selectable effort levels (low, medium, high, xhigh, max). Effort is the main control: output token spend runs roughly 8x from low to max, and Artificial Analysis measures a 407-Elo spread in task quality across that range, so the same model behaves like several different price and capability tiers. API pricing is $5 per million input tokens and $25 per million output, with cache writes at $6.25 and cache hits at $0.50. It leads the Intelligence Index at 60.7 and tops the Coding Agent Index, scoring 89 percent on Terminal-Bench v2.1 at max effort and 53 percent on Humanity's Last Exam.
Why: It is the current number one on the independent Artificial Analysis Intelligence Index, and it got there while costing less per task than the model it displaced: $2.03 average per index task against Fable 5's $2.75. The effort dial is the reason to pick it over a fixed-tier model, because one integration covers cheap high-volume calls and expensive long-horizon agent runs.
Freemium Best Frontier Model Overall Visit
The Efficiency Revolution: Frontier Intelligence at 1/100th the Cost
Added Feb 5, 2026
DeepSeek is the architect of the 'DeepSeek movement,' a fundamental shift in AI development that prioritizes extreme efficiency over raw compute. Founded by High-Flyer Quant, they proved that architectural innovations like Multi-head Latent Attention (MLA) and DeepSeekMoE could match the performance of $100B models like GPT-4o and Claude 3.5 while costing 95% less to train and run. Their ecosystem includes the flagship DeepSeek-V3, the reasoning-heavy DeepSeek-R1, and the state-of-the-art DeepSeek-VL2 for high-fidelity OCR and vision tasks. DeepSeek is committed to the open-source community, regularly releasing model weights and technical papers that have democratized frontier-level AI for developers globally.
Why: DeepSeek changed the game by proving that 'expensive' doesn't always mean 'better.' We picked it because it's the first model family to offer true frontier-level reasoning (R1), general intelligence (V3), and advanced vision/OCR (VL2) with an open-weight philosophy and an API price point that makes proprietary models look obsolete.
Freemium Best for Cost-Efficiency Visit
The 'Next DeepSeek' Movement: o1-Level Reasoning at 1/100th the Cost
Added Jan 31, 2026
Kimi k1.5 is a multimodal large language model from Moonshot AI, specifically engineered for high-fidelity technical reasoning and long-context processing. It is a key player in the 'DeepSeek movement,' matching the reasoning performance of frontier models like GPT-5.2 Codex and Claude 4.5 while remaining significantly more cost-effective. It features a massive 2 million token context window and joint text-vision reasoning, making it ideal for complex coding, mathematical proofs, and large-scale document analysis. The model is built using advanced Reinforcement Learning (RL) to achieve deep 'Chain-of-Thought' capabilities.
Why: Kimi k1.5 is the first model to prove that o1-level reasoning is achievable through efficient, open-weight architectures. We selected it because it consistently matches or exceeds Claude 4.5 in technical benchmarks (AIME, MATH-500) while offering a 2M context window and a significantly lower API price point, making frontier intelligence accessible to everyone.
Freemium Best for Technical Reasoning Visit
The Open Vision-Reasoner: SOTA Multimodal Performance
Added Jan 31, 2026
Qwen 2.5-VL is Alibaba's state-of-the-art open-weight multimodal model, designed to bridge the gap between open source and proprietary vision-language models. It features advanced 'NaViVi' (Native Dynamic Resolution) architecture, allowing it to process images of any resolution and videos of any length with extreme precision. It excels at complex visual reasoning, document understanding (OCR), and real-time video analysis, matching or exceeding GPT-4o in many multimodal benchmarks while remaining fully open for the community to build upon.
Why: We added Qwen 2.5-VL to the Open Frontier movement because it is currently the highest-performing open-weight vision model. It proves that open source can lead in multimodal reasoning, especially for tasks requiring high-resolution OCR and long-form video understanding.
Free Best for Open Vision Reasoning Visit
Meta's Open Multimodal Standard
Added Jan 31, 2026
Llama 3.2 Vision is Meta's first open-weight multimodal model family, bringing high-fidelity image reasoning to the Llama ecosystem. It integrates vision and text into a unified transformer architecture, enabling it to understand images, charts, and diagrams with the same ease as text. Available in 11B and 90B versions, it is designed for efficiency and edge deployment, making it the industry standard for developers building open multimodal applications that require deep reasoning and broad community support.
Why: We included Llama 3.2 Vision because it is the most widely supported open multimodal model in the world. Its integration into almost every AI tool and framework makes it the 'default' choice for open-weight vision reasoning.
Free Best for Open Ecosystem Support Visit
The Open Vision Frontier: 124B Multimodal Power
Added Jan 31, 2026
Pixtral Large is Mistral AI's flagship 124B parameter multimodal model, designed to compete directly with GPT-4o and Claude 3.5 Sonnet. Built on the Mistral Large 2 foundation, it features a native vision encoder that allows it to reason across text and images with extreme precision. It excels at complex diagram understanding, mathematical reasoning with visual context, and high-fidelity image captioning. Pixtral Large is released under the Mistral Research License, allowing developers to explore frontier-level vision-language capabilities with open weights.
Why: We added Pixtral Large because it represents the peak of European open-weight AI. It is one of the few open models that truly matches the visual reasoning depth of the top proprietary models, making it essential for the Open Frontier movement.
Freemium Best for Complex Visual Reasoning Visit
The Open-Source Vision Giant: 78B Multimodal Leader
Added Jan 31, 2026
InternVL 2.5 is a world-class open-source multimodal large language model (MLLM) that consistently tops the leaderboards for open-weight vision reasoning. It features a powerful 78B parameter architecture with a specialized vision-language alignment that excels at OCR, document understanding, and complex visual Q&A. It is designed to bridge the gap between open models and GPT-4V, offering exceptional performance across a wide range of multimodal benchmarks while remaining fully open for community development.
Why: We included InternVL 2.5 because it is a consistent leaderboard champion. It often outperforms much larger models in visual reasoning and OCR, making it a critical tool for developers who need GPT-4 level vision without the proprietary lock-in.
Free Best for Leaderboard-Topping Vision Visit
Google's fast, capable multimodal model from I/O 2026
Added May 19, 2026
Gemini 3.5 Flash is a mid-tier multimodal model announced at Google I/O on May 19, 2026. It delivers strong reasoning, coding, and long-context performance at lower latency and cost than Ultra-tier models, with native support for text, images, audio, and video inputs.
Why: Gemini 3.5 Flash hits a practical sweet spot for developers and creators who need more capability than entry-level models but do not require the full cost of an Ultra model. Its native multimodal design makes it especially useful for mixed-media tasks.
Freemium Best for Fast Multimodality Visit
Google's unified multimodal generation model
Added May 19, 2026
Gemini Omni is a single Google model announced at I/O 2026 that can generate and reason across text, images, video, and audio from unified prompts. It is designed to reduce the need for separate modality-specific models by handling generation and understanding in one architecture.
Why: Gemini Omni represents Google's push toward a single model for all media types. For teams building multimodal products, it simplifies architecture by replacing multiple specialized endpoints with one interface.
Freemium Best for Unified Generation Visit
DeepSeek's open-weight model with permanent pricing
Added May 31, 2026
DeepSeek V4-Pro is a high-performance language model from DeepSeek. The model itself was released on April 24, 2026, and permanent pricing was announced on May 31, 2026. It offers strong reasoning and coding performance at a competitive price point, with open weights available for local deployment.
Why: DeepSeek V4-Pro stands out for combining frontier-level performance with transparent, permanent pricing and open weights. It is a practical choice for teams that want to self-host or avoid unpredictable API costs.
Freemium Best for Predictable Pricing Visit
StepFun's 198B MoE vision-language model
Added May 29, 2026
StepFun Step 3.7 Flash is a 198-billion-parameter mixture-of-experts vision-language model released on May 28-29, 2026. It supports text, image, and video understanding with a focus on efficient inference and strong multimodal reasoning.
Why: Step 3.7 Flash offers a competitive Chinese-frontier multimodal model with an MoE architecture that balances capability and inference cost. It is a useful option for vision-language applications and for teams exploring alternatives to US models.
Freemium Best for Efficient VLM Visit
The frontier model for complex reasoning and software architecture
Added Feb 6, 2026
OpenAI's most advanced model to date, featuring a 2M context window and specialized training for complex multi-step reasoning. It excels at architectural planning, deep research, and autonomous code generation.
Why: GPT-5.3 Codex is the 'World's Smartest Planner.' While other AI tools are good at chatting, this one is built for solving huge, difficult problems like planning how a whole software system should work. It has a massive memory (2 million words) so it never loses track of the big picture.
Paid Best for Reasoning Visit
Native multimodal intelligence with a 10M context window
Added Feb 5, 2026
Google's most powerful multimodal model, capable of processing hours of video, thousands of lines of code, or massive document sets in a single prompt. Features native audio/video understanding.
Why: Gemini 3 Ultra offers an unbeatable 10M token context window, allowing it to process entire project histories, hours of video, or massive codebases in a single prompt. Its native multimodal intelligence makes it the only model capable of 'seeing' and 'hearing' complex data sets with the same level of depth as it reads text, providing a unique advantage for large-scale data analysis.
Paid Best for Context Visit
One multimodal model for text, vision, audio, and video reasoning
Added May 3, 2026
Nemotron 3 Nano Omni is NVIDIA's compact-but-capable multimodal stack for agentic workflows: one family of endpoints that accept text, images, audio, or video (depending on route) and return text answers, useful as the 'perception and reasoning' layer for assistants that must read screens, documents, calls, or clips without chaining four different specialist models. Optimized for efficiency at scale; exposed on fal.ai as separate text, vision, audio, and video reasoning endpoints built on the same foundation.
Why: If your product roadmap says 'agents that see and hear the world,' Omni is built for that integration story, fewer moving parts than bolting Whisper + CLIP + LLM together by hand.
Paid Best for Agents Visit
Alibaba's strongest vision-language model
Added Jan 29, 2024
Qwen-VL-Max is a high-performance vision-language model from Alibaba, capable of understanding images, charts, and documents, and answering questions about them. It is available through the Qwen API and Tongyi Qianwen apps.
Freemium Best for Vision-Language Visit
Frontier Opus model with higher-resolution vision and xhigh effort
Added Apr 16, 2026
Anthropic's frontier Opus-tier model announced on April 16, 2026, with substantial gains on the hardest coding tasks, higher-resolution vision input, a new xhigh effort level, file-system-based memory recall across sessions, and the Task Budgets public beta. Available across the Anthropic API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry.
Why: Opus 4.7 introduced the xhigh effort level and file-system memory recall, making it a notable step between Opus 4.6 and Opus 4.8.
Enterprise Best for Hard Coding Tasks Visit
Limited-availability Mythos-class model without Fable 5 safety classifiers
Added Jun 9, 2026
Anthropic's Mythos-class model announced on June 9, 2026, shares the same capabilities as Claude Fable 5 without the safety classifiers. It is offered only in limited availability to approved customers through Anthropic's Project Glasswing program.
Why: Mythos 5 is a notable limited-availability variant of the Mythos-class tier, distinct from the generally available Fable 5.
Enterprise Best for Controlled Research Visit
High-volume DeepSeek inference with a 1M-token context window
Added Apr 24, 2026
DeepSeek V4-Flash is the efficient sibling of V4-Pro, offering a 1M-token context window and configurable thinking modes at a fraction of the API cost. It is optimized for high-throughput chat, classification, bulk extraction, and agentic coding workloads, with a July 2026 update that boosted agent and coding benchmarks.
Why: V4-Flash delivers the same 1M context and thinking modes as V4-Pro at roughly one-third the API cost, making it the practical default for most production workloads.
Freemium Best for High-Volume APIs Visit
The open-weight reasoning model that sparked the efficiency revolution
Added Jan 20, 2025
DeepSeek R1 is a 671B-parameter open-weight reasoning model that matches o1-class performance on math, code, and logic benchmarks through reinforcement learning on verifiable tasks. It exposes chain-of-thought reasoning and is available as MIT-licensed local weights and via API, with the R1-0528 update in May 2025 further improving math and code reasoning.
Why: R1 proved that open-weight models can match proprietary reasoning systems at a fraction of the cost, making it a landmark for reproducible AI research.
Freemium Best for Open Reasoning Visit
The 128K-context MoE flagship that introduced sparse attention
Added Dec 1, 2025
DeepSeek V3.2 is a 128K-context mixture-of-experts model that unified thinking and non-thinking modes in the V3 line and introduced DeepSeek Sparse Attention. It served as the December 2025 flagship before V4 and remains available for self-hosting and as a historical comparison point.
Why: V3.2 introduced DeepSeek Sparse Attention and unified thinking modes, making it the architectural bridge that enabled the later 1M-context V4 family.
Freemium Best for Long-Context MoE Visit
Multimodal coding and visual-reasoning agent model
Added Jun 1, 2026
GLM-5V-Turbo is a vision-language variant of the GLM-5 family, built for multimodal coding, visual reasoning, and image-plus-text agent workflows. It offers a 200K context window and 128K maximum output.
Why: GLM-5V-Turbo is the GLM family's main vision agent, letting coding and agent workflows reason over images and screenshots in the same long context.
Paid Best for Multimodal Coding Visit
Vision-language model for visual reasoning and UI replication
Added Dec 1, 2025
GLM-4.6V is a 2025 vision-language model in the GLM family, offering visual reasoning, tool calling, and frontend code replication. It supports a 128K context window and a 32K maximum output.
Why: GLM-4.6V is the practical vision tier for turning screenshots and images into working code or structured analysis.
Paid Best for Visual Reasoning Visit
Document parsing model for PDF and image OCR
Added Oct 1, 2025
GLM-OCR is a specialized GLM model for extracting structured Markdown from PDFs and images. It supports a 65K context window and is optimized for layout-aware document parsing rather than open-ended chat.
Why: GLM-OCR fills a clear gap in the GLM family by turning scanned documents and PDFs into structured, usable text with layout awareness.
Paid Best for Document Parsing Visit
Free vision model for image understanding and document snapshots
Added Apr 1, 2026
GLM-4V-Flash is a free-tier vision model in the GLM family, offering zero-cost image understanding and document snapshot analysis with a 16K context window.
Why: GLM-4V-Flash is the entry-level vision option for GLM, letting users test multimodal document understanding before upgrading to paid vision tiers.
Free Best for Free Vision Visit
Google's long-context multimodal flagship with up to 2M tokens
Added Feb 15, 2024
Gemini 1.5 Pro is a mid-2024 multimodal model that handles text, images, audio, and video with a 1M-token context window (extendable to 2M in limited preview). It powers complex document analysis, code understanding, and video QA in Google AI Studio and the Gemini API.
Freemium Best for Long Context Visit
Fast, cost-efficient multimodal model with a 1M context window
Added May 21, 2024
Gemini 1.5 Flash is a lightweight, speed-optimized variant of Gemini 1.5 Pro. It keeps the same 1M-token context window and multimodal input support while offering much lower latency and cost, making it ideal for high-volume agents and summarization.
Freemium Best for Fast Multimodal Tasks Visit
Google's low-latency agentic model with native tool use
Added Dec 11, 2024
Gemini 2.0 Flash is a late-2024 general-purpose model optimized for agentic workflows, native tool use, and fast multimodal output. It supports text, image, audio, and video input and is the default model for many Gemini API applications.
Freemium Best for Agentic Apps Visit
Google's high-performance reasoning model with advanced coding
Added Mar 25, 2025
Gemini 2.5 Pro is an early-2025 flagship model focused on complex reasoning, advanced coding, and detailed multimodal understanding. It builds on Gemini 2.0 with improved instruction following and is positioned for high-stakes enterprise and research tasks.
Freemium Best for Complex Reasoning Visit
Google's open multimodal model for research and developers
Added Mar 12, 2025
Gemma 3 is an open-weights family of multimodal models from Google, ranging from 1B to 27B parameters. It supports text and image input, a 128K context window, and is released under a permissive license for research and commercial use.
Free Best for Open Multimodal Visit
xAI's long-context flagship with a 1M-token window
Added Apr 1, 2026
Grok 4.3 is xAI's general-purpose frontier model released in April 2026. It pairs strong reasoning and coding with a 1-million-token context window, making it practical for analyzing long documents and large codebases in a single pass.
Why: Grok 4.3 is the sweet spot in xAI's lineup for anyone who needs a frontier model with a very large context window at a lower price than Grok 4.5.
Paid Best for Long-Context Work Visit
xAI's 2M-context beta model with multi-agent capabilities
Added Feb 1, 2026
Grok 4.20 is a beta model from xAI released in February 2026. It features a 2-million-token context window and multi-agent architecture, targeting complex reasoning and long-horizon workflows that benefit from coordinated sub-agents.
Why: Grok 4.20 remains notable as xAI's first multi-agent beta model with a 2M context window, even though newer 4.3/4.5 models now offer flagship alternatives.
Paid Best for Multi-Agent Beta Work Visit
Tencent's Mamba-powered deep-thinking reasoning model
Added Mar 21, 2025
Hybrid Mamba-Transformer MoE reasoning model released March 2025, built on Hunyuan TurboS with 52 billion active parameters and a 256K context window. It focuses compute on reinforcement-learning post-training and scores strongly on math, coding, and graduate-level reasoning tasks.
Why: One of the first ultra-large Mamba-Transformer MoE reasoning models, offering strong benchmark scores and a 256K context window.
Freemium Best for Reasoning Visit
The deep-thinking variant of Hunyuan 2.0
Added Sep 15, 2025
Open-weight reasoning variant of Hunyuan 2.0 with a 131K context window, designed for complex problem-solving, math, and long-context reasoning workflows.
Why: Hunyuan 2.0's reasoning mode for tasks that benefit from longer thought chains.
Freemium Best for Reasoning Visit
Moonshot's open-weight multimodal generalist with agent swarms
Added Jan 27, 2026
Kimi K2.5 is Moonshot AI's open-weight multimodal model released January 27, 2026. It supports text, image, and video input, thinking and non-thinking modes, and a 256K context window, with strong performance on agent, coding, and vision tasks. Moonshot announced the kimi-k2.5 API will be retired on August 31, 2026, so production workloads should plan a migration path to K2.6 or K3.
Why: Kimi K2.5 was Moonshot's first widely available open-weight multimodal generalist and remains a notable reference point for the K2 family before K2.6 and K3 arrived.
Freemium Best for Open Multimodal Agents Visit
Moonshot's open-weight multimodal successor with long-context coding stability
Added Apr 21, 2026
Kimi K2.6 is Moonshot AI's open-weight multimodal model released April 21, 2026. It supports text, image, and video input, thinking and non-thinking modes, and a 256K context window, with improved long-context coding stability and agent-task performance.
Why: Kimi K2.6 improves on K2.5 with stronger long-context coding and is a practical open-weight alternative for teams that want multimodal agents without the cost of closed frontier models.
Freemium Best for Long-Context Coding Visit
Meta's open-weight flagship with native multimodal reasoning
Added Apr 5, 2025
Llama 4 Maverick is Meta's flagship open-weight model, released in April 2025 as part of the Llama 4 family. It uses a Mixture-of-Experts architecture with 17B active parameters and around 400B total parameters, natively understands text and images, and supports a 1M-token context window.
Why: Maverick is the top open-weight model Meta actually ships today, with strong multimodal reasoning and a practical API ecosystem, making it the default choice for open Llama deployments.
Free Best for Open Multimodal Reasoning Visit
Long-context, efficient open multimodal model for edge and single-GPU use
Added Apr 5, 2025
Llama 4 Scout is Meta's efficient Llama 4 variant, released in April 2025. It is a Mixture-of-Experts model with 17B active parameters and 109B total parameters across 16 experts, natively multimodal for text and images, and supports an industry-leading 10M-token context window.
Why: Scout is notable for its extreme 10M-token context window and efficient single-GPU deployment, making it the standout open model for very long documents and memory-heavy applications.
Free Best for Long Context Visit
Mistral's flagship open-weight multimodal frontier model
Added Apr 15, 2026
A 675B-parameter sparse mixture-of-experts model with 41B active parameters and a 262K context window, released under Apache 2.0. It handles text and vision tasks, supports strong multilingual performance, and is designed for both research and enterprise deployment.
Why: Mistral Large 3 is one of the most capable permissive open-weight models available, offering frontier performance with the deployment flexibility of Apache 2.0 licensing.
Freemium Best for Open-Weight Frontier Visit
Unified open-source small model for chat, reasoning, vision, and coding
Added May 1, 2026
A 119B-parameter MoE model with 6B active parameters and a 256K context window, released under Apache 2.0. It unifies instruct, reasoning, multimodal, and agentic coding capabilities in a single efficient model with configurable reasoning effort.
Why: Small 4 packs flagship-class reasoning, vision, and coding into a single open-source model that is efficient enough for high-throughput and local deployments.
Freemium Best for Efficient Open Multimodal Visit
Layout-aware document parsing that goes beyond OCR
Added Jun 1, 2025
A general-purpose document parsing model that overcomes traditional OCR limitations by understanding complex page layouts. It extracts structured text, tables as Markdown or LaTeX, bounding boxes, and semantic classes from unstructured PDFs and images.
Why: Turns messy documents into clean structured data, improving downstream RAG and agent pipelines with layout-aware extraction.
Free Best for Document Intelligence Visit
Multimodal 4B safety model for text and image moderation
Added Jun 4, 2026
A 4B-parameter multimodal, multilingual small language model designed as a robust content-safety moderator. It supports standard taxonomy safety classification and custom-policy enforcement with reasoning traces for text and image inputs.
Why: A compact, open safety model that can enforce both standard and custom content policies with reasoning traces for safer deployments.
Free Best for Content Safety Visit
Open foundation model for humanoid robot reasoning and control
Added Jun 4, 2026
NVIDIA's open foundation model for humanoid robot reasoning and control, combining an Eagle-based vision-language backbone with a diffusion transformer (DiT) action head for language-conditioned manipulation across diverse embodiments.
Why: NVIDIA's open contribution to humanoid robot foundation models, enabling language-conditioned manipulation across diverse robot embodiments.
Free Best for Humanoid Robotics Visit
Microsoft's unified AI model family from Build 2026
Added Jul 7, 2026
Microsoft announced a family of MAI-branded models at Build 2026, including MAI-Thinking-1 for reasoning, MAI-Image-2.5 for image generation and editing, MAI-Voice-2 for expressive text-to-speech, MAI-Transcribe-1.5 for speech-to-text, MAI-Code-1-Flash for coding in GitHub Copilot, and Scout as a workplace personal agent. They integrate tightly with Microsoft 365, Azure, and GitHub.
Why: The MAI family gives Microsoft a cohesive, enterprise-ready AI stack. For organizations already using Microsoft services, these models reduce friction by running inside familiar tools rather than requiring separate platforms.
Enterprise Best for Microsoft Ecosystem Visit
Meta's closed-weight agentic model, and its first paid model API
Added Aug 4, 2026
Muse Spark 1.1 is Meta Superintelligence Labs' multimodal reasoning model, released 9 July 2026. It has a 1M-token context window, accepts text, images, video and PDFs, and is built for agentic work: tool use, computer use, coding, and multi-agent orchestration. It is free to use in the Meta AI app and at meta.ai in Thinking mode, and available to developers through the new Meta Model API at $1.25 per million input tokens and $4.25 per million output, with $20 in starting credits. Unlike the Llama family it succeeds in practice, Muse Spark is closed-weight.
Why: This is the release where Meta stopped giving models away. Muse Spark is closed, metered and sold through Meta's own API, and it lands at rank 15 on the independent index while undercutting comparable models on price. Worth tracking for that reason alone if your stack assumed Meta meant open weights.
Freemium Best Value for Agentic Multimodal Work Visit
Thinking Machines' 975B Apache-2.0 model that takes text, images and audio natively
Added Aug 4, 2026
Inkling is the first open-weights model from Thinking Machines Lab, released 15 July 2026 under Apache 2.0. It is a 975B-parameter mixture of experts with 41B active per token, trained from scratch on 45 trillion tokens of text, images, audio and video, with a context window up to 1M tokens. The architecture uses 256 routed experts plus 2 shared experts per layer with 6 routed experts active per token, a sigmoid router, and interleaved sliding-window and global attention at a 5:1 ratio. It accepts text, image and audio input natively and returns text. Weights are on Hugging Face in both the original format and an NVFP4 checkpoint for Blackwell hardware. A distilled Inkling-Small followed on 31 July.
Why: It is the first roughly trillion-parameter open-weights model that takes audio and images natively rather than through a bolted-on encoder, and Apache 2.0 means the weights can be used commercially without asking anyone. Thinking Machines is candid that this is not the strongest model available but a base worth customising, which is a more honest pitch than most open releases make.
Free Best Open-Weight Multimodal Base Visit