No official DeepSWE mini-swe-agent row for this identity on the public board we fetched.
nvidia/nemotron-3-ultra-550b-a55b:freeMoE 550B-A55B from the OpenRouter id and the official repo name NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16. Free list price on that slug. Continuum free hosted id matches.
HF config.json (fetched 19 Aug 2026): NemotronHForCausalLM, hidden 8192, 512 routed experts, 22 experts/token, 1 shared expert. License on the Hub card is OpenMDW 1.1 (license field other, license_name openmdw-1.1).
Use Ultra when you want a free hosted Nemotron. AA Intelligence Index 38; vals hero index 27.39%. No official DeepSWE row.
No official DeepSWE mini-swe-agent row for this identity on the public board we fetched.
vals.ai Nemotron 3 Ultra card, fetched 19 Aug 2026: hero Index 27.39% ±0.95, 32 min 28 s latency. A 4 Jun 2026 Updates paragraph still prints 43.99% among open-weight on the Index plus TaxEval v2 73.10%, CorpFin v2 65.46%, Finance Agent v2 37.53%, MedCode 38.62%, GPQA Diamond 86.11%, MMLU Pro 85.76%, LegalBench 82.07%, LiveCodeBench 85.98%, Terminal Bench 2.1 50.94%. Those named percents are Updates prose, not the 0.0% Accuracy Rankings bars (omitted), and 43.99% is not today’s hero.
Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.
Copied from AA or vals HTML we opened. 0.0% placeholder bars are omitted. Do not average these into the hero tiles, and do not invent CI, $/task, or steps for Artificial Analysis.
| Bench | Printed | Note | Source | URL | As of |
|---|---|---|---|---|---|
| AA list pair (Reasoning tile) | $0.60 / $2.75 per 1M | Printed AA pricing sentence. OpenRouter free slug is $0 / $0. | Artificial Analysis Nemotron 3 Ultra | source | 2026-08-19 |
| AA throughput | 128.7 tok/s | Printed Speed row. Not a coding score. | Artificial Analysis Nemotron 3 Ultra | source | 2026-08-19 |
| AA Index eval cost | $534.18 | Printed “it cost $534.18 to evaluate Nemotron 3 Ultra 550B A55B (Reasoning) on the Intelligence Index.” | Artificial Analysis Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates (4 Jun 2026) open-weight Index | 43.99% | Printed in Updates prose. Current hero is 27.39% ±0.95. Do not swap them. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates TaxEval v2 | 73.10% | Updates prose, among open-weight. Not a 0.0% bar. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates CorpFin v2 | 65.46% | Updates prose. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates Finance Agent v2 | 37.53% | Updates prose. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates MedCode | 38.62% | Updates prose. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates GPQA Diamond | 86.11% | Updates prose. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates MMLU Pro | 85.76% | Updates prose. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates LegalBench | 82.07% | Updates prose. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates LiveCodeBench | 85.98% | Updates prose. Not DeepSWE. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
| vals Updates Terminal Bench 2.1 | 50.94% | Updates prose. Not a hero chip. | vals.ai Nemotron 3 Ultra | source | 2026-08-19 |
Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.
Community quants: Community quants are community.
This is the model repo, not Download for Mac. Downloads and likes from the Hugging Face API on 2026-08-19.
usedStorage and the sum of .safetensors sizes can differ (LFS pointers, non-shard files, duplicate copies). Cited separately. Source: recursive tree API, 2026-08-19.
| Path | Bytes |
|---|---|
model-00001-of-00225.safetensors | 4,983,097,968 |
model-00002-of-00225.safetensors | 4,991,252,496 |
model-00003-of-00225.safetensors | 4,991,252,600 |
model-00004-of-00225.safetensors | 4,991,252,600 |
model-00005-of-00225.safetensors | 4,920,152,112 |
model-00006-of-00225.safetensors | 4,995,491,168 |
model-00007-of-00225.safetensors | 4,991,252,600 |
model-00008-of-00225.safetensors | 4,991,252,600 |
model-00009-of-00225.safetensors | 4,991,252,600 |
model-00010-of-00225.safetensors | 4,987,305,680 |
model-00011-of-00225.safetensors | 4,991,252,464 |
model-00012-of-00225.safetensors | 4,991,252,600 |
Showing the first 12 of 225 safetensor shards. The tree also holds 16 non-shard files, none of them listed here. The snapshot was capped at 12 shards when the tree API was read on 2026-08-19, so the remaining 213 are on the Files tab, not on this page. Totals above are the full tree API sum, not this slice. Byte sizes are Hub tree API size fields, never invented.
Continuum free allowlist includes nvidia/nemotron-3-ultra-550b-a55b:free (19 Aug 2026).
nvidia/nemotron-3-ultra-550b-a55b:free
click to select
We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.
OpenRouter free slug nvidia/nemotron-3-ultra-550b-a55b:free, 19 Aug 2026: $0 / $0 list, 1,000,000 / 65,536, created 1780551208. No cache fields. Free is not unlimited.
Artificial Analysis page title is “Nemotron 3 Ultra 550B A55B (Reasoning)”. That page prints $0.60 / $2.75 per 1M, 128.7 tok/s, and “it cost $534.18 to evaluate … on the Intelligence Index.” Those cents are the paid Reasoning tile, not the free slug.
Artificial Analysis Intelligence Index 38, fetched 19 Aug 2026, on the Reasoning tile. Same page prints $0.60 / $2.75, 128.7 tok/s, and $534.18 to run the Index. The AA tile is the published integer 38. Do not copy AA’s paid pair onto the free Continuum id.
| Claim | Source | As of |
|---|---|---|
| OR free slug $0 / $0, 1,000,000 / 65,536, created 1780551208, hugging_face_id NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | OpenRouter catalog | 2026-08-19 |
| 2.69T week tokens on the free Ultra slug | OpenRouter rankings This Week | 2026-08-19 |
| Hosted free id | GET /v1/chat/hosted/models/public | 2026-08-19 |
| HF safetensors.total 560,524,578,816; usedStorage 2,242,147,855,302; 225 shards; likes 328; created 2026-06-03T14:50:04Z; license_name openmdw-1.1 | Hugging Face API NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | 2026-08-19 |
| AA title Reasoning; Index 38; $0.60 / $2.75; 128.7 tok/s; $534.18 to evaluate Index | Artificial Analysis Nemotron 3 Ultra | 2026-08-19 |
| vals hero 27.39% ±0.95; 32 min 28 s; Updates 4 Jun 2026 open-weight Index 43.99% plus named prose benches | vals.ai Nemotron 3 Ultra | 2026-08-19 |
| OpenRouter id nvidia/nemotron-3-ultra-550b-a55b:free, context 1,000,000 | OpenRouter /api/v1/models | 2026-08-19 |
| HF downloads 443,962 | Hugging Face API nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | 2026-08-19 |
| 225 safetensor shards; shard bytes 1,121,055,890,600; tree files 241 | Hugging Face tree API nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | 2026-08-19 |
| Spaces API returned 1; official collection NVIDIA Nemotron v3 (4 items, 360 upvotes) | Hugging Face Spaces / collections API | 2026-08-19 |
No other card in this cluster shares a published DeepSWE official row we can put next to this one. We will not compare on vendor-blog numbers.
Continue through NVIDIA's model family with Nemotron 3.5 Lightning.
The official DeepSWE board updated 13 Aug 2026 did not list Nemotron 3 Ultra. Omitting is the rule.
No. Continuum hosts the :free slug at $0 / $0. AA’s $0.60 / $2.75 is the paid Reasoning page. Cite both; do not invoice free at paid.
27.39% ±0.95 is today’s vals hero. 43.99% is a 4 Jun 2026 Updates sentence about open-weight rank. The chip stays the hero.
Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.
Verified against the live public allowlist on 2026-08-19. Base URL is https://continuumcode.ai/v1. Keys are cont_sk_ from Settings, Account, Inference API. Personal keys need Plus or above.
curl https://continuumcode.ai/v1/chat/completions \
-H "Authorization: Bearer $CONTINUUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"nvidia/nemotron-3-ultra-550b-a55b:free","messages":[{"role":"user","content":"Review this diff."}]}'