- bench
- DeepSWE
- version
- v1.1
- split
- public 113 tasks
- harness
- mini-swe-agent
- effort
- max
- n
- 113
- metric
- Pass@1
- value
- 69%
- ci
- ±5%
- $/task
- $4.65
- tokens
- 81k out
- steps
- 98
kimi-k3Official Hub repo moonshotai/Kimi-K3. License field is other, license_name kimi-k3. Read the repo LICENSE before you ship it. That is not MIT.
HF config.json (raw 200, fetched 19 Aug 2026): KimiK3ForConditionalGeneration, text stack KimiLinearForCausalLM, hidden 7168, first_k_dense_replace 1.
Vendor README (raw 200, 19 Aug 2026): 2.8T total / 104B activated, 93 layers (1 dense + 69 KDA + 24 Gated MLA), hidden 7168, 96 heads, 896 experts / 16 selected, 1M context. Those are vendor claims. Hub safetensors.total is 2,779,931,837,184 stored tensors. Do not average the two.
Official quickstart: thinking stays on. reasoning_effort is low / high / max, default max. max_completion_tokens default 131,072, max 1,048,576. Temperature / top_p fixed at 1.0 / 0.95.
Use K3 when you want an open Moonshot checkpoint Continuum already hosts and you accept the kimi-k3 license. DeepSWE official (mini-swe-agent, 113 tasks): 69% Pass@1 ±5% at max, $4.65 per task, 81k output, 98 steps.
Against Opus (74% ±4, $11.84) the intervals overlap. Against Flash (53% ±4, $0.10) you are buying pass rate, not cents. Against Fable (70% ±4, $21.63) you are buying open weights at a lower DeepSWE dollar.
vals.ai Kimi K3 card, fetched 19 Aug 2026: Index 57.81% ±1.06, $6.468 per Index test, 69 min 34 s latency. Those are Index-page extras, not DeepSWE, and not hero chips. The per-bench bars on that HTML were placeholders (0.0%) in this fetch. We will not invent SWE-bench percents from them.
Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.
Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.
Community quants: Any GGUF or AWQ you find under other orgs is community, not official.
Official collection Kimi K3 (2 items, 111 upvotes) is Hub popularity, not a quality ranking.
Vendor 2.8T / 104B is README copy. The hero tensor count on this page is Hub safetensors.total.
Hub API listed 39 spaces on this repo (19 Aug 2026). That is Hub popularity, not a quality ranking.
This is the model repo, not Download for Mac. Downloads and likes from the Hugging Face API on 2026-08-19.
usedStorage and the sum of .safetensors sizes can differ (LFS pointers, non-shard files, duplicate copies). Cited separately. Source: recursive tree API, 2026-08-19.
| Path | Bytes |
|---|---|
model-00001-of-000096.safetensors | 2,341,216,112 |
model-00002-of-000096.safetensors | 16,990,911,504 |
model-00003-of-000096.safetensors | 16,990,911,504 |
model-00004-of-000096.safetensors | 16,567,501,776 |
model-00005-of-000096.safetensors | 16,990,911,504 |
model-00006-of-000096.safetensors | 16,990,911,504 |
model-00007-of-000096.safetensors | 16,990,911,504 |
model-00008-of-000096.safetensors | 16,567,501,776 |
model-00009-of-000096.safetensors | 16,990,911,504 |
model-00010-of-000096.safetensors | 16,990,911,504 |
model-00011-of-000096.safetensors | 16,990,916,912 |
model-00012-of-000096.safetensors | 16,567,507,176 |
Showing the first 12 of 96 safetensor shards. The tree also holds 22 non-shard files, none of them listed here. The snapshot was capped at 12 shards when the tree API was read on 2026-08-19, so the remaining 84 are on the Files tab, not on this page. Totals above are the full tree API sum, not this slice. Byte sizes are Hub tree API size fields, never invented.
Continuum hosts kimi-k3 on the live public allowlist fetched 19 Aug 2026.
First-party Moonshot API and OpenRouter moonshotai/kimi-k3 are the other hosts we will name. We will not invent a 15-host table.
kimi-k3
click to select
We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.
DeepSWE official: $4.65 per task at max (81k out, 98 steps). That is the hero economics number.
Official Kimi K3 pricing page, fetched 19 Aug 2026: cache-hit / cache-miss / output $0.30 / $3.00 / $15.00 per million tokens; context 1,048,576. Taxes excluded at checkout.
OpenRouter catalog 19 Aug 2026: $3 / $15 per million, cache read $0.30 /M, context 1,048,576, hugging_face_id moonshotai/Kimi-K3. Continuum rate card matches $3 / $15. Official miss/out and the catalog list are the same pair; cache-hit $0.30 is the official third column.
Artificial Analysis Intelligence Index 60, fetched 19 Aug 2026. AA technical specs on the same page: 2800B / 104B active, 38.3 tok/s, about $0.84 per AA task at max, list $3 / $15. The vendor README’s 2.8T / 104B is a different citation of the same scale; do not average the two. The AA tile is the published integer 60 plus its harness label. Do not average AA cost with DeepSWE $4.65/task.
| Claim | Source | As of |
|---|---|---|
| DeepSWE 69% ±5% at max | DeepSWE official board | 2026-08-13 |
| Official K3 cache-hit / miss / out $0.30 / $3.00 / $15.00 per 1M; context 1,048,576 | Kimi K3 pricing | 2026-08-19 |
| OR catalog $3 / $15, cache read $0.30/M, hugging_face_id moonshotai/Kimi-K3 | OpenRouter catalog | 2026-08-19 |
| HF safetensors.total 2,779,931,837,184; usedStorage 1,561,018,243,668; 96 shards; likes 10,848; downloads 2,289,863; license_name kimi-k3; dtypes F32 11,122,432 / BF16 57,179,884,544 / U8 2,722,740,830,208; 39 spaces | Hugging Face API moonshotai/Kimi-K3 | 2026-08-19 |
| AA Index 60; 38.3 tok/s; 2800B / 104B; $0.84/AA task; $3 / $15 | Artificial Analysis Kimi K3 | 2026-08-19 |
| vals Index 57.81% ±1.06; $6.468/test; 69 min 34 s | vals.ai Kimi K3 | 2026-08-19 |
| Vendor README 2.8T / 104B activated; config KimiK3ForConditionalGeneration, hidden 7168 | HF README + config.json Kimi-K3 | 2026-08-19 |
| Hosted id kimi-k3 | GET /v1/chat/hosted/models/public | 2026-08-19 |
| 1.32T week tokens (−9% WoW) | OpenRouter rankings This Week | 2026-08-19 |
| OpenRouter id moonshotai/kimi-k3, context 1,048,576 | OpenRouter /api/v1/models | 2026-08-19 |
| HF downloads 2,289,863 | Hugging Face API moonshotai/Kimi-K3 | 2026-08-19 |
| 96 safetensor shards; shard bytes 1,560,936,091,448; tree files 118 | Hugging Face tree API moonshotai/Kimi-K3 | 2026-08-19 |
| Spaces API returned 39; official collection Kimi K3 (2 items, 111 upvotes) | Hugging Face Spaces / collections API | 2026-08-19 |
Same official DeepSWE harness (mini-swe-agent, public 113 tasks), fetched 2026-08-19. Different effort labels are the lab's own setting on that board, shown here rather than normalized.
Continue through Moonshot's model family with Kimi K2.7 Code.
The Hub license field is other and license_name is kimi-k3. Read the official LICENSE in the repo. This page will not relabel it MIT.
No. 2.8T / 104B is vendor README copy. Hub safetensors.total is 2,779,931,837,184. We cite both and do not average them.
Official quickstart: thinking stays on. reasoning_effort is low / high / max, default max.
Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.
Verified against the live public allowlist on 2026-08-19. Base URL is https://continuumcode.ai/v1. Keys are cont_sk_ from Settings, Account, Inference API. Personal keys need Plus or above.
curl https://continuumcode.ai/v1/chat/completions \
-H "Authorization: Bearer $CONTINUUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"kimi-k3","messages":[{"role":"user","content":"Review this diff."}]}'