Nvidia: Nemotron 3 Ultra
Open-weight agentic reasoning model with 55B active parameters, 1M-token (~977K) context window, and vendor-reported SWE-bench Verified score of 71.9%. Priced at $0.50 per million input tokens vs $2.00 for Claude Sonnet 5. Ideal for coding agents and enterprise deployments needing long-context, low-cost inference.
Specifications
| Attribute | Value |
|---|---|
| Lab | Nvidia |
| Tags | Intelligent Agentic Open Weight Long Context Reasoning |
| Release Date | 2026-06 |
| Context Window | 1,000,000 tokens |
| Input Price / 1M | $0.50 |
| Output Price / 1M | $2.50 |
| Input Modalities | Text |
| Output Modalities | Text |
Strengths
- 71.9% on SWE-bench Verified (vendor-reported) for coding agent tasks
- ~977K token context window at $0.50/M input tokens
- 55B active parameters out of 550B total for efficient MoE inference
- Open-weight release under OpenMDW License v1.1 with weights and data
- 5.9× higher throughput than GLM-5.1-754B-A40B per vendor claims
Weaknesses
- Independent benchmark verification still pending
- Requires 8× GB200 or 16× H100 minimum hardware
- English-dominant performance; limited multilingual coverage
- LMRank 8.5/10 trails Claude Sonnet 5 (9.4) and GPT-5.5 (9.4)
Best For
- agent orchestration and long-running coding agents
- deep research with 1M-token context window
- enterprise agentic AI with open-weight customization
- multilingual reasoning across 10 languages
In Depth: Nemotron 3 Ultra
Benchmark Performance
NVIDIA: Nemotron 3 Ultra (#37), while Claude Fable 5 leads at 9.9/10. On vendor-reported SWE-bench Verified, Nemotron 3 Ultra reaches 71.9% a strong result for an open-weight model at 12.5× lower input cost than GPT-5.5 Pro.
NVIDIA: Nemotron 3 Ultra agentic reasoning model reports a SWE-bench Verified score of 71.9% and GPQA 87.0% (vendor-reported via Nemo Evaluator SDK). On long-context retrieval (RULER @ 1M) it achieves 94.7%, and PinchBench 90.0%. These benchmarks are self-reported and pending independent verification. The model uses Multi-Token Prediction layers for native speculative decoding, boosting throughput 5.9× over GLM-5.1-754B-A40B.
Pricing & Value
NVIDIA: Nemotron 3 Ultra pricing starts at $0.50 per million input tokens and $2.50 per million output tokens. For comparison, Claude Sonnet 5 charges $2.00 input and $10.00 output 4× more for both directions.
At $0.50 input / $2.50 output per million tokens, Nemotron 3 Ultra costs 40× less input than Claude Opus 4.8 ($15.00/$75.00) while delivering an LMRank score of 8.5 vs 9.7. That is $0.059 per LMRank point per million input tokens versus $1.55 for Opus 4.8 a 96% savings per point. Claude Sonnet 5 ($2.00/$10.00) costs 4× more input and scores 9.4, just 0.9 points higher.
Who Should Use This
NVIDIA: Nemotron 3 Ultra agentic reasoning model suits three personas. Agent engineers gain low-cost, long-context inference for coding agents. Enterprise AI teams get fully open weights for customization. Deep research teams leverage the ~977K context window.
- Agent engineers: deploy long-running coding agents at 40× lower input cost than Claude Opus 4.8, but you must verify benchmarks independently.
- Enterprise AI teams: customize open-weight weights and data under OpenMDW License v1.1, though hardware requires 8× GB200 or 16× H100 minimum.
- Deep research teams: exploit ~977K context window for document analysis, with caveat that English-dominant performance may limit multilingual use.
- Avoid if: you need independently verified benchmark scores or prefer a higher LMRank model like Claude Fable 5 (9.9) for production-critical tasks.
Release & Version History
NVIDIA: Nemotron 3 Ultra released in June 2026, arriving as the final model in the Nemotron 3 family. It offers a 550B total parameter (55B active) MoE architecture with a ~977K token context window, the largest in the series.
June 4, 2026: Nemotron 3 Ultra launches as the final and largest model in the Nemotron 3 family. It succeeds the Nemotron 3 series with a 550B total parameter (55B active) Mixture-of-Experts architecture, hybrid LatentMoE-Mamba-2+MoE+Attention design, and Multi-Token Prediction layers. Pre-trained in NVFP4 precision on 20 trillion text tokens, it extends context from the previous 128K to ~977K tokens. Released under the OpenMDW License v1.1 with weights, data, and recipes on Hugging Face.