Tested and written up.
Added Jun 4, 2026
A 30B total / 3B active parameter hybrid Mamba-2 + Transformer MoE language model built for efficient on-device and edge agentic tasks. It features a 1M-token context window, reasoning ON/OFF modes with configurable thinking budgets, and up to 4× faster throughput than its predecessor.
Why: The smallest open-weight member of the Nemotron 3 family, giving teams frontier-style reasoning and tool-use without data-center hardware.
Added Jun 4, 2026
A 120B total / 12B active parameter hybrid Mamba-Transformer MoE language model with LatentMoE, multi-token prediction, and native NVFP4 pretraining. Optimized for complex multi-agent applications with a 1M-token context window and up to 5× higher throughput than the previous Nemotron Super.
Why: Fills the gap between Nano and Ultra with a strong efficiency-to-accuracy ratio for agentic orchestration and latency-sensitive serving.
Added Jun 4, 2026
A 4B-parameter multimodal, multilingual small language model designed as a robust content-safety moderator. It supports standard taxonomy safety classification and custom-policy enforcement with reasoning traces for text and image inputs.
Why: A compact, open safety model that can enforce both standard and custom content policies with reasoning traces for safer deployments.
Added Jun 4, 2026
NVIDIA's open foundation model for humanoid robot reasoning and control, combining an Eagle-based vision-language backbone with a diffusion transformer (DiT) action head for language-conditioned manipulation across diverse embodiments.
Why: NVIDIA's open contribution to humanoid robot foundation models, enabling language-conditioned manipulation across diverse robot embodiments.
Added Jun 1, 2026
GLM-5-Turbo is a tuned variant of the GLM-5 series that prioritizes lower latency and efficient sequential execution. It shares the 200K context window and 128K output ceiling of the GLM-5 family and is positioned for agentic workflows that need many quick steps.
Why: GLM-5-Turbo is the practical speed layer for the GLM-5 family, trading a small amount of peak capability for noticeably faster multi-step agent execution.
Added Jun 1, 2026
GLM-5V-Turbo is a vision-language variant of the GLM-5 family, built for multimodal coding, visual reasoning, and image-plus-text agent workflows. It offers a 200K context window and 128K maximum output.
Why: GLM-5V-Turbo is the GLM family's main vision agent, letting coding and agent workflows reason over images and screenshots in the same long context.
Added Jun 1, 2026
MiniMax Speech 2.8 HD generates ultra-realistic, expressive speech with sound tags, supporting 40 languages, 7 emotions, and specified dialects for high-fidelity voice applications.
Why: Speech 2.8 HD is the current quality-tier MiniMax voice model, replacing the earlier Speech 2.6 / Speech-02 series.
Added Jun 1, 2026
MiniMax Speech 2.8 Turbo balances speed and naturalness, supporting 40 languages, 7 emotions, and specified dialects for real-time, low-latency voice synthesis.
Why: Speech 2.8 Turbo is the current speed-tier MiniMax voice model, distinct from the HD quality variant.
Added Jun 1, 2026
MiniMax Music 3.0 generates music from text and reference inputs with improved intent understanding, elevated sound quality, and more humanized vocals compared to the previous Music 2.0 generation.
Why: Music 3.0 is the current MiniMax music generation model, replacing the legacy Music 2.0 entry already in the directory.
Added Jun 1, 2026
An 8B-class physical-AI omni-model (8B reasoner + 8B generator) optimized for efficient inference on workstation-grade NVIDIA hardware such as the RTX PRO 6000. It unifies vision reasoning, world generation, and action prediction for robotics and physical AI prototyping.
Why: Brings Cosmos 3 physical-AI capabilities to smaller hardware so labs and individual developers can experiment without a data center.
Added Jun 1, 2026
A 4B-class physical-AI omni-model (2B reasoner + 2B generator) optimized for real-time robotic policy and visual reasoning at the edge. It unifies world generation, vision reasoning, and action prediction in a compact form factor for embedded deployment.
Why: The smallest Cosmos 3 variant, built for real-time robotic perception and policy where latency and power matter most.