Models/NVIDIA/Nemotron 3 Ultra

Nemotron 3 Ultra

Free-lane Ultra. Continuum lists the same free id. 550B total / 55B active is in the catalog id, not a guess.

01

Identity

1,000,000Context in
65,536Context out
2026-06-04 (OpenRouter created 1780551208)Released
OpenWeights
Canonical name
Nemotron 3 Ultra
Aliases
nemotron-3-ultra, nvidia/nemotron-3-ultra-550b-a55b:free
Continuum hosted id
nvidia/nemotron-3-ultra-550b-a55b:free
OpenRouter slug
nvidia/nemotron-3-ultra-550b-a55b:free
Hugging Face
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
Modalities
text
Open weights Hosted on Continuum Vendor lab

MoE 550B-A55B from the OpenRouter id and the official repo name NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16. Free list price on that slug. Continuum free hosted id matches.

HF config.json (fetched 19 Aug 2026): NemotronHForCausalLM, hidden 8192, 512 routed experts, 22 experts/token, 1 shared expert. License on the Hub card is OpenMDW 1.1 (license field other, license_name openmdw-1.1).

02

Should I use this for coding agents

Use Ultra when you want a free hosted Nemotron. AA Intelligence Index 38; vals hero index 27.39%. No official DeepSWE row.

DeepSWE

No official DeepSWE mini-swe-agent row for this identity on the public board we fetched.

omitted, not invented as of 2026-08-13 source
Artificial Analysispublished integer
38
bench
Artificial Analysis
version
Intelligence Index
harness
published integer
metric
Intelligence Index
value
38
independent as of 2026-08-19 source
Vals Indexvals.ai card
27.39%
bench
Vals Index
version
hero index
harness
vals.ai card
metric
Vals Index
value
27.39%
independent as of 2026-08-19 source

vals.ai Nemotron 3 Ultra card, fetched 19 Aug 2026: hero Index 27.39% ±0.95, 32 min 28 s latency. A 4 Jun 2026 Updates paragraph still prints 43.99% among open-weight on the Index plus TaxEval v2 73.10%, CorpFin v2 65.46%, Finance Agent v2 37.53%, MedCode 38.62%, GPQA Diamond 86.11%, MMLU Pro 85.76%, LegalBench 82.07%, LiveCodeBench 85.98%, Terminal Bench 2.1 50.94%. Those named percents are Updates prose, not the 0.0% Accuracy Rankings bars (omitted), and 43.99% is not today’s hero.

Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.

Named benches, printed only

Copied from AA or vals HTML we opened. 0.0% placeholder bars are omitted. Do not average these into the hero tiles, and do not invent CI, $/task, or steps for Artificial Analysis.

BenchPrintedNoteSourceURLAs of
AA list pair (Reasoning tile) $0.60 / $2.75 per 1M Printed AA pricing sentence. OpenRouter free slug is $0 / $0. Artificial Analysis Nemotron 3 Ultra source 2026-08-19
AA throughput 128.7 tok/s Printed Speed row. Not a coding score. Artificial Analysis Nemotron 3 Ultra source 2026-08-19
AA Index eval cost $534.18 Printed “it cost $534.18 to evaluate Nemotron 3 Ultra 550B A55B (Reasoning) on the Intelligence Index.” Artificial Analysis Nemotron 3 Ultra source 2026-08-19
vals Updates (4 Jun 2026) open-weight Index 43.99% Printed in Updates prose. Current hero is 27.39% ±0.95. Do not swap them. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates TaxEval v2 73.10% Updates prose, among open-weight. Not a 0.0% bar. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates CorpFin v2 65.46% Updates prose. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates Finance Agent v2 37.53% Updates prose. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates MedCode 38.62% Updates prose. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates GPQA Diamond 86.11% Updates prose. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates MMLU Pro 85.76% Updates prose. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates LegalBench 82.07% Updates prose. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates LiveCodeBench 85.98% Updates prose. Not DeepSWE. vals.ai Nemotron 3 Ultra source 2026-08-19
vals Updates Terminal Bench 2.1 50.94% Updates prose. Not a hero chip. vals.ai Nemotron 3 Ultra source 2026-08-19
03

When not to use it

Trust this before you buy

  • No official DeepSWE row. Free-lane rate limits are not a paid SLO. Do not plan capacity on a :free id. OpenRouter list on this slug is $0 / $0; AA’s page is the paid Reasoning tile at $0.60 / $2.75. Do not invoice a free lane at AA’s pair.
  • vals hero is 27.39% ±0.95 today. A 4 Jun 2026 Updates paragraph still prints 43.99% among open-weight on the Index. Those are two dates; the hero chip stays 27.39%.
  • OpenMDW 1.1, not MIT. No official GGUF on this BF16 repo. Hub usedStorage (2,242,147,855,302) is about 2× the tree shard-byte sum (1,121,055,890,600), cited separately, not averaged.
  • Sibling: Nemotron 3.5 Lightning is the other free NVIDIA slug on the week board, not this card. AA title is “Nemotron 3 Ultra 550B A55B (Reasoning)”.

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.

04

Artifact and Hugging Face downloads

Artifact

SPDX / license
OpenMDW 1.1 (HF license other; license_name openmdw-1.1)
Total params
HF safetensors.total: 560,524,578,816 stored tensors. Catalog / repo name: 550B. Do not average those two.
Activated params
55B (catalog id A55B and official repo name). Not a Hub safetensors field.
Architecture
NemotronHForCausalLM · hidden 8192 · 512 routed experts · 22 experts/tok · 1 shared · BF16 official repo
Native precision
HF tensors BF16 + F32. Official repo suffix BF16
Files
225 safetensor shards (model-00001-of-00225 …)
Repo size
HF API usedStorage 2,242,147,855,302 bytes (2,088 GiB)
HF created
2026-06-03T14:50:04Z
Chat template
Jinja present (chat_template.jinja).
Paper
No arXiv tag on the Hugging Face API payload we fetched.
Official HF repo
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
HF downloads
443,962
HF likes
328

Community quants: Community quants are community.

Official weights

This is the model repo, not Download for Mac. Downloads and likes from the Hugging Face API on 2026-08-19.

Hub file tree

Safetensor shards
225
Shard bytes
1,121,055,890,600 bytes
Other file bytes
22,557,926 bytes
Tree file count
241
Hub usedStorage
2,242,147,855,302 bytes

usedStorage and the sum of .safetensors sizes can differ (LFS pointers, non-shard files, duplicate copies). Cited separately. Source: recursive tree API, 2026-08-19.

PathBytes
model-00001-of-00225.safetensors4,983,097,968
model-00002-of-00225.safetensors4,991,252,496
model-00003-of-00225.safetensors4,991,252,600
model-00004-of-00225.safetensors4,991,252,600
model-00005-of-00225.safetensors4,920,152,112
model-00006-of-00225.safetensors4,995,491,168
model-00007-of-00225.safetensors4,991,252,600
model-00008-of-00225.safetensors4,991,252,600
model-00009-of-00225.safetensors4,991,252,600
model-00010-of-00225.safetensors4,987,305,680
model-00011-of-00225.safetensors4,991,252,464
model-00012-of-00225.safetensors4,991,252,600

Showing the first 12 of 225 safetensor shards. The tree also holds 16 non-shard files, none of them listed here. The snapshot was capped at 12 shards when the tree API was read on 2026-08-19, so the remaining 213 are on the Files tab, not on this page. Totals above are the full tree API sum, not this slice. Byte sizes are Hub tree API size fields, never invented.

05

Continuum serving

Continuum free allowlist includes nvidia/nemotron-3-ultra-550b-a55b:free (19 Aug 2026).

Continuum hosted id nvidia/nemotron-3-ultra-550b-a55b:free

click to select

We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.

06

Economics

OpenRouter free slug nvidia/nemotron-3-ultra-550b-a55b:free, 19 Aug 2026: $0 / $0 list, 1,000,000 / 65,536, created 1780551208. No cache fields. Free is not unlimited.

Artificial Analysis page title is “Nemotron 3 Ultra 550B A55B (Reasoning)”. That page prints $0.60 / $2.75 per 1M, 128.7 tok/s, and “it cost $534.18 to evaluate … on the Intelligence Index.” Those cents are the paid Reasoning tile, not the free slug.

Price this model Opens the pricing calculator preloaded with Nemotron 3 Ultra.

Artificial Analysis Intelligence Index 38, fetched 19 Aug 2026, on the Reasoning tile. Same page prints $0.60 / $2.75, 128.7 tok/s, and $534.18 to run the Index. The AA tile is the published integer 38. Do not copy AA’s paid pair onto the free Continuum id.

07

Provenance

ClaimSourceAs of
OR free slug $0 / $0, 1,000,000 / 65,536, created 1780551208, hugging_face_id NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16OpenRouter catalog2026-08-19
2.69T week tokens on the free Ultra slugOpenRouter rankings This Week2026-08-19
Hosted free idGET /v1/chat/hosted/models/public2026-08-19
HF safetensors.total 560,524,578,816; usedStorage 2,242,147,855,302; 225 shards; likes 328; created 2026-06-03T14:50:04Z; license_name openmdw-1.1Hugging Face API NVIDIA-Nemotron-3-Ultra-550B-A55B-BF162026-08-19
AA title Reasoning; Index 38; $0.60 / $2.75; 128.7 tok/s; $534.18 to evaluate IndexArtificial Analysis Nemotron 3 Ultra2026-08-19
vals hero 27.39% ±0.95; 32 min 28 s; Updates 4 Jun 2026 open-weight Index 43.99% plus named prose benchesvals.ai Nemotron 3 Ultra2026-08-19
OpenRouter id nvidia/nemotron-3-ultra-550b-a55b:free, context 1,000,000OpenRouter /api/v1/models2026-08-19
HF downloads 443,962Hugging Face API nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF162026-08-19
225 safetensor shards; shard bytes 1,121,055,890,600; tree files 241Hugging Face tree API nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF162026-08-19
Spaces API returned 1; official collection NVIDIA Nemotron v3 (4 items, 360 upvotes)Hugging Face Spaces / collections API2026-08-19
08

Compare, FAQ, and Get Plus

No other card in this cluster shares a published DeepSWE official row we can put next to this one. We will not compare on vendor-blog numbers.

Continue through NVIDIA's model family with Nemotron 3.5 Lightning.

FAQ

Why no DeepSWE chip?

The official DeepSWE board updated 13 Aug 2026 did not list Nemotron 3 Ultra. Omitting is the rule.

Is the Continuum id the AA $0.60 tile?

No. Continuum hosts the :free slug at $0 / $0. AA’s $0.60 / $2.75 is the paid Reasoning page. Cite both; do not invoice free at paid.

Why 27.39% and 43.99%?

27.39% ±0.95 is today’s vals hero. 43.99% is a 4 Jun 2026 Updates sentence about open-weight rank. The chip stays the hero.

Why does the lab blog disagree with DeepSWE or Scale?

Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.

OpenAI-compatible call

Verified against the live public allowlist on 2026-08-19. Base URL is https://continuumcode.ai/v1. Keys are cont_sk_ from Settings, Account, Inference API. Personal keys need Plus or above.

curl https://continuumcode.ai/v1/chat/completions \
  -H "Authorization: Bearer $CONTINUUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"nvidia/nemotron-3-ultra-550b-a55b:free","messages":[{"role":"user","content":"Review this diff."}]}'

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.