Nvidia: Nemotron 3 Ultra

by Nvidia Intelligent Agentic Open Weight Long Context Reasoning

Open-weight agentic reasoning model with 55B active parameters, 1M-token (~977K) context window, and vendor-reported SWE-bench Verified score of 71.9%. Priced at $0.50 per million input tokens vs $2.00 for Claude Sonnet 5. Ideal for coding agents and enterprise deployments needing long-context, low-cost inference.

Choose a model to compare against Nemotron 3 Ultra

Specifications

Specifications for Nemotron 3 Ultra
AttributeValue
Lab Nvidia
Tags Intelligent Agentic Open Weight Long Context Reasoning
Release Date 2026-06
Context Window 1,000,000 tokens
Input Price / 1M $0.50
Output Price / 1M $2.50
Input Modalities Text
Output Modalities Text

Strengths

  • 71.9% on SWE-bench Verified (vendor-reported) for coding agent tasks
  • ~977K token context window at $0.50/M input tokens
  • 55B active parameters out of 550B total for efficient MoE inference
  • Open-weight release under OpenMDW License v1.1 with weights and data
  • 5.9× higher throughput than GLM-5.1-754B-A40B per vendor claims

Weaknesses

  • Independent benchmark verification still pending
  • Requires 8× GB200 or 16× H100 minimum hardware
  • English-dominant performance; limited multilingual coverage
  • LMRank 8.5/10 trails Claude Sonnet 5 (9.4) and GPT-5.5 (9.4)

Best For

  • agent orchestration and long-running coding agents
  • deep research with 1M-token context window
  • enterprise agentic AI with open-weight customization
  • multilingual reasoning across 10 languages

In Depth: Nemotron 3 Ultra

Benchmark Performance

NVIDIA: Nemotron 3 Ultra (#37), while Claude Fable 5 leads at 9.9/10. On vendor-reported SWE-bench Verified, Nemotron 3 Ultra reaches 71.9% a strong result for an open-weight model at 12.5× lower input cost than GPT-5.5 Pro.

NVIDIA: Nemotron 3 Ultra agentic reasoning model reports a SWE-bench Verified score of 71.9% and GPQA 87.0% (vendor-reported via Nemo Evaluator SDK). On long-context retrieval (RULER @ 1M) it achieves 94.7%, and PinchBench 90.0%. These benchmarks are self-reported and pending independent verification. The model uses Multi-Token Prediction layers for native speculative decoding, boosting throughput 5.9× over GLM-5.1-754B-A40B.

Pricing & Value

NVIDIA: Nemotron 3 Ultra pricing starts at $0.50 per million input tokens and $2.50 per million output tokens. For comparison, Claude Sonnet 5 charges $2.00 input and $10.00 output 4× more for both directions.

At $0.50 input / $2.50 output per million tokens, Nemotron 3 Ultra costs 40× less input than Claude Opus 4.8 ($15.00/$75.00) while delivering an LMRank score of 8.5 vs 9.7. That is $0.059 per LMRank point per million input tokens versus $1.55 for Opus 4.8 a 96% savings per point. Claude Sonnet 5 ($2.00/$10.00) costs 4× more input and scores 9.4, just 0.9 points higher.

Who Should Use This

NVIDIA: Nemotron 3 Ultra agentic reasoning model suits three personas. Agent engineers gain low-cost, long-context inference for coding agents. Enterprise AI teams get fully open weights for customization. Deep research teams leverage the ~977K context window.

  • Agent engineers: deploy long-running coding agents at 40× lower input cost than Claude Opus 4.8, but you must verify benchmarks independently.
  • Enterprise AI teams: customize open-weight weights and data under OpenMDW License v1.1, though hardware requires 8× GB200 or 16× H100 minimum.
  • Deep research teams: exploit ~977K context window for document analysis, with caveat that English-dominant performance may limit multilingual use.
  • Avoid if: you need independently verified benchmark scores or prefer a higher LMRank model like Claude Fable 5 (9.9) for production-critical tasks.

Release & Version History

NVIDIA: Nemotron 3 Ultra released in June 2026, arriving as the final model in the Nemotron 3 family. It offers a 550B total parameter (55B active) MoE architecture with a ~977K token context window, the largest in the series.

June 4, 2026: Nemotron 3 Ultra launches as the final and largest model in the Nemotron 3 family. It succeeds the Nemotron 3 series with a 550B total parameter (55B active) Mixture-of-Experts architecture, hybrid LatentMoE-Mamba-2+MoE+Attention design, and Multi-Token Prediction layers. Pre-trained in NVFP4 precision on 20 trillion text tokens, it extends context from the previous 128K to ~977K tokens. Released under the OpenMDW License v1.1 with weights, data, and recipes on Hugging Face.

Sources & Further Reading

Related Models