Back
All NVIDIA models
LLMS • CURATED • UPDATED JUN 4, 2026

NVIDIA Nemotron 3 Super

120B open-weight hybrid MoE for efficient multi-agent reasoning

A 120B total / 12B active parameter hybrid Mamba-Transformer MoE language model with LatentMoE, multi-token prediction, and native NVFP4 pretraining. Optimized for complex multi-agent applications with a 1M-token context window and up to 5× higher throughput than the previous Nemotron Super.

Pricing Free
Platforms api, local
API Yes
Open Source Yes
Modalities LLMs
Best For Best for Multi-Agent Efficiency
Date Added 2026-06-04

Fills the gap between Nano and Ultra with a strong efficiency-to-accuracy ratio for agentic orchestration and latency-sensitive serving.

Website Hugging Face
Claude Fable 5 GPT-6 Astra NotebookLM Claude Opus 5 GPT-5.6 Sol
Free Completely free
📚

Open Models: What They Cannot Do

The benchmark gap has largely closed: open models now score within a point of the closed frontier on...

Hardware For Local AI Models: What The Memory Shortage Changed

Capacity decides whether a model runs. Bandwidth decides how fast it talks, and it is the number mos...

How To Run AI Models Locally: Hardware, Quantisation And Engines

Running a frontier open model locally is a memory problem before it is a speed problem, and a bandwi...

Uncensored AI Models: What Abliteration Actually Does To Them

Abliteration removes a model's ability to refuse by deleting one direction from its activations. It ...

Which LLMs Actually Deliver in 2026?

Comprehensive comparison of the best large language models in 2026 including ChatGPT, Claude, Gemini...

View NVIDIA Nemotron 3 Super Alternatives (2026) →

Compare NVIDIA Nemotron 3 Super with 5+ similar llms AI tools.

Q

Is NVIDIA Nemotron 3 Super free?

A

Yes, NVIDIA Nemotron 3 Super is completely free to use.

Q

Does NVIDIA Nemotron 3 Super have an API?

A

Yes, NVIDIA Nemotron 3 Super offers an API for programmatic integration.

Q

What is NVIDIA Nemotron 3 Super best for?

A

NVIDIA Nemotron 3 Super is best for Multi-Agent Efficiency. A 120B total / 12B active parameter hybrid Mamba-Transformer MoE language model with LatentMoE, multi-token prediction, and native NVFP4 pretraining. Fills the gap between Nano and Ultra with a strong efficiency-to-accuracy ratio for agentic orchestration and latency-sensitive serving.

Q

What platforms does NVIDIA Nemotron 3 Super support?

A

NVIDIA Nemotron 3 Super supports api, local.

Q

Is NVIDIA Nemotron 3 Super open source?

A

Yes, NVIDIA Nemotron 3 Super is open source. You can access the source code, contribute, and deploy it on your own infrastructure.

Q

How do I use NVIDIA Nemotron 3 Super?

A

NVIDIA Nemotron 3 Super is a large language model for text generation, analysis, and conversation. Use the API for programmatic access. Enter prompts or questions to get responses.

🏷️

Work on NVIDIA Nemotron 3 Super? You're hand-reviewed in our directory. Add this badge to your site — it links back to this profile.

Featured on CuratedAI