Skip to main content

NVIDIA models

NVIDIA's Nemotron models are open-weight releases tuned for helpfulness and agentic tasks, built on top of leading open architectures and optimized for NVIDIA hardware.

All 10 NVIDIA models

Nemotron 3.5 Lightning (free)Open weight

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

1M ctx0 cr/M outAugust 2026
Nemotron 3.5 LightningOpen weight

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...

262K ctx2.2 cr/M outAugust 2026
Nemotron 3.5 Content SafetyOpen weight

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

131K ctx2.2 cr/M outJune 2026
Nemotron 3.5 Content Safety (free)Open weight

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

128K ctx0 cr/M outJune 2026
Nemotron 3 Ultra (free)Open weight

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

1M ctx0 cr/M outJune 2026
Nemotron 3 UltraOpen weight

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

256K ctx34.4 cr/M outJune 2026
Nemotron 3 Nano Omni (free)Fast and light

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

256K ctx0 cr/M outApril 2026
Nemotron 3 SuperOpen weight

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

262K ctx4.4 cr/M outMarch 2026
Nemotron 3 Super (free)Open weight

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

262K ctx0 cr/M outMarch 2026
Nemotron 3 Nano 30B A3BFast and light

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

262K ctx2.2 cr/M outDecember 2025

Pricing at a glance

ModelContextInput cr/MOutput cr/MReleased
Nemotron 3.5 Lightning (free)1M00August 2026
Nemotron 3.5 Lightning262K0.882.2August 2026
Nemotron 3.5 Content Safety131K2.22.2June 2026
Nemotron 3.5 Content Safety (free)128K00June 2026
Nemotron 3 Ultra (free)1M00June 2026
Nemotron 3 Ultra256K6.934.4June 2026
Nemotron 3 Nano Omni (free)256K00April 2026
Nemotron 3 Super262K0.934.4March 2026
Nemotron 3 Super (free)262K00March 2026
Nemotron 3 Nano 30B A3B262K0.552.2December 2025

Other providers

Explore Zeplik

NVIDIA Models - Pricing, Context and Capabilities | Zeplik Chat