NVIDIA models
NVIDIA's Nemotron models are open-weight releases tuned for helpfulness and agentic tasks, built on top of leading open architectures and optimized for NVIDIA hardware.
All 10 NVIDIA models
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
Pricing at a glance
| Model | Context | Input cr/M | Output cr/M | Released |
|---|---|---|---|---|
| Nemotron 3.5 Lightning (free) | 1M | 0 | 0 | August 2026 |
| Nemotron 3.5 Lightning | 262K | 0.88 | 2.2 | August 2026 |
| Nemotron 3.5 Content Safety | 131K | 2.2 | 2.2 | June 2026 |
| Nemotron 3.5 Content Safety (free) | 128K | 0 | 0 | June 2026 |
| Nemotron 3 Ultra (free) | 1M | 0 | 0 | June 2026 |
| Nemotron 3 Ultra | 256K | 6.9 | 34.4 | June 2026 |
| Nemotron 3 Nano Omni (free) | 256K | 0 | 0 | April 2026 |
| Nemotron 3 Super | 262K | 0.93 | 4.4 | March 2026 |
| Nemotron 3 Super (free) | 262K | 0 | 0 | March 2026 |
| Nemotron 3 Nano 30B A3B | 262K | 0.55 | 2.2 | December 2025 |
Other providers
- OpenAI (90)
- Qwen (52)
- Google (43)
- Anthropic (27)
- Mistral (20)
- Z.ai (16)
- DeepSeek (16)
- MiniMax (11)
- All models
Explore Zeplik
- 465 AI skillsReady-to-run expert methods that run on any NVIDIA model, or any other model you pick.
- 997 integrationsConnect Gmail, Slack, GitHub, Notion and hundreds more so the assistant can work in your apps.
- Simple pricingPlus, Pro and Max plans with a monthly credit allowance that works across every model.