Kimi K3

Moonshot’s current coding identity. #13 on OpenRouter This Week at 1.32T (−9% week-over-week, not share). Open weights, hosted on Continuum. DeepSWE 69% ±5% at max is the hero row: a band, not a point.

01

Identity

1,048,576Context in
Not published on OpenRouterContext out
2026-07-16 (OpenRouter created)Released
OpenWeights
Canonical name
Kimi K3
Aliases
kimi-k3, moonshotai/kimi-k3
Continuum hosted id
kimi-k3
OpenRouter slug
moonshotai/kimi-k3
Hugging Face
moonshotai/Kimi-K3
Modalities
text, image, video
Open weights Hosted on Continuum Vendor lab

Official Hub repo moonshotai/Kimi-K3. License field is other, license_name kimi-k3. Read the repo LICENSE before you ship it. That is not MIT.

HF config.json (raw 200, fetched 19 Aug 2026): KimiK3ForConditionalGeneration, text stack KimiLinearForCausalLM, hidden 7168, first_k_dense_replace 1.

Vendor README (raw 200, 19 Aug 2026): 2.8T total / 104B activated, 93 layers (1 dense + 69 KDA + 24 Gated MLA), hidden 7168, 96 heads, 896 experts / 16 selected, 1M context. Those are vendor claims. Hub safetensors.total is 2,779,931,837,184 stored tensors. Do not average the two.

Official quickstart: thinking stays on. reasoning_effort is low / high / max, default max. max_completion_tokens default 131,072, max 1,048,576. Temperature / top_p fixed at 1.0 / 0.95.

02

Should I use this for coding agents

Use K3 when you want an open Moonshot checkpoint Continuum already hosts and you accept the kimi-k3 license. DeepSWE official (mini-swe-agent, 113 tasks): 69% Pass@1 ±5% at max, $4.65 per task, 81k output, 98 steps.

Against Opus (74% ±4, $11.84) the intervals overlap. Against Flash (53% ±4, $0.10) you are buying pass rate, not cents. Against Fable (70% ±4, $21.63) you are buying open weights at a lower DeepSWE dollar.

DeepSWEmini-swe-agent
69%±5%
bench
DeepSWE
version
v1.1
split
public 113 tasks
harness
mini-swe-agent
effort
max
n
113
metric
Pass@1
value
69%
ci
±5%
$/task
$4.65
tokens
81k out
steps
98
independent as of 2026-08-13 source
Artificial Analysispublished integer
60
bench
Artificial Analysis
version
Intelligence Index
harness
published integer
metric
Intelligence Index
value
60
independent as of 2026-08-19 source
Vals Indexvals.ai card
57.81%
bench
Vals Index
version
hero index
harness
vals.ai card
metric
Vals Index
value
57.81%
independent as of 2026-08-19 source

vals.ai Kimi K3 card, fetched 19 Aug 2026: Index 57.81% ±1.06, $6.468 per Index test, 69 min 34 s latency. Those are Index-page extras, not DeepSWE, and not hero chips. The per-bench bars on that HTML were placeholders (0.0%) in this fetch. We will not invent SWE-bench percents from them.

Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.

03

When not to use it

Trust this before you buy

  • ±5% is wide. Do not treat 69% as a point estimate against Opus 74% ±4. The bands overlap.
  • Thinking is always on. Default effort is max. If you have not tried a cheaper DeepSWE row (Luna $0.61, Flash $0.10), do not start here for mechanical volume.
  • License is kimi-k3, not MIT. There is no official GGUF on this repo. Community quants are community.
  • Sibling: Kimi K2.7 Code is the other Moonshot card. Flash is the cheap open DeepSWE alternative; Opus / Sol if you need the top closed rows.

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.

04

Artifact and Hugging Face downloads

Artifact

SPDX / license
other / license_name kimi-k3 (read official LICENSE). Not MIT.
Total params
Vendor README: 2.8T total. HF safetensors.total: 2,779,931,837,184 stored tensors. Do not average those two.
Activated params
104B (vendor README, fetched 19 Aug 2026). Not a Hub safetensors field.
Architecture
KimiK3ForConditionalGeneration · text KimiLinearForCausalLM · hidden 7168 · first_k_dense_replace 1. Vendor README: 93 layers (1 dense + 69 KDA + 24 Gated MLA), 96 heads, 896 experts / 16 selected
Native precision
Hub safetensors.parameters (API 19 Aug 2026): F32 11,122,432 · BF16 57,179,884,544 · U8 2,722,740,830,208. Those are parameter counts by dtype, not a per-shard table.
Files
96 safetensor shards (model-00001-of-00096 …)
Repo size
HF API usedStorage 1,561,018,243,668 bytes. Tree API shard-byte sum 1,560,936,091,448.
HF created
2026-06-13T06:42:57Z
Sampling
Official quickstart: thinking always on. reasoning_effort low | high | max, default max. temp 1.0 / top_p 0.95. max_completion_tokens default 131072, max 1048576.
Chat template
Present on the Hub card we fetched.
Official HF repo
moonshotai/Kimi-K3
HF downloads
2,289,863
HF likes
10,848

Community quants: Any GGUF or AWQ you find under other orgs is community, not official.

Official collection Kimi K3 (2 items, 111 upvotes) is Hub popularity, not a quality ranking.

Vendor 2.8T / 104B is README copy. The hero tensor count on this page is Hub safetensors.total.

Hub API listed 39 spaces on this repo (19 Aug 2026). That is Hub popularity, not a quality ranking.

Official weights

This is the model repo, not Download for Mac. Downloads and likes from the Hugging Face API on 2026-08-19.

Hub file tree

Safetensor shards
96
Shard bytes
1,560,936,091,448 bytes
Other file bytes
62,892,942 bytes
Tree file count
118
Hub usedStorage
1,561,018,243,668 bytes

usedStorage and the sum of .safetensors sizes can differ (LFS pointers, non-shard files, duplicate copies). Cited separately. Source: recursive tree API, 2026-08-19.

PathBytes
model-00001-of-000096.safetensors2,341,216,112
model-00002-of-000096.safetensors16,990,911,504
model-00003-of-000096.safetensors16,990,911,504
model-00004-of-000096.safetensors16,567,501,776
model-00005-of-000096.safetensors16,990,911,504
model-00006-of-000096.safetensors16,990,911,504
model-00007-of-000096.safetensors16,990,911,504
model-00008-of-000096.safetensors16,567,501,776
model-00009-of-000096.safetensors16,990,911,504
model-00010-of-000096.safetensors16,990,911,504
model-00011-of-000096.safetensors16,990,916,912
model-00012-of-000096.safetensors16,567,507,176

Showing the first 12 of 96 safetensor shards. The tree also holds 22 non-shard files, none of them listed here. The snapshot was capped at 12 shards when the tree API was read on 2026-08-19, so the remaining 84 are on the Files tab, not on this page. Totals above are the full tree API sum, not this slice. Byte sizes are Hub tree API size fields, never invented.

05

Continuum serving

Continuum hosts kimi-k3 on the live public allowlist fetched 19 Aug 2026.

First-party Moonshot API and OpenRouter moonshotai/kimi-k3 are the other hosts we will name. We will not invent a 15-host table.

Continuum hosted id kimi-k3

click to select

We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.

06

Economics

DeepSWE official: $4.65 per task at max (81k out, 98 steps). That is the hero economics number.

Official Kimi K3 pricing page, fetched 19 Aug 2026: cache-hit / cache-miss / output $0.30 / $3.00 / $15.00 per million tokens; context 1,048,576. Taxes excluded at checkout.

OpenRouter catalog 19 Aug 2026: $3 / $15 per million, cache read $0.30 /M, context 1,048,576, hugging_face_id moonshotai/Kimi-K3. Continuum rate card matches $3 / $15. Official miss/out and the catalog list are the same pair; cache-hit $0.30 is the official third column.

Price this model Opens the pricing calculator preloaded with Kimi K3.

Artificial Analysis Intelligence Index 60, fetched 19 Aug 2026. AA technical specs on the same page: 2800B / 104B active, 38.3 tok/s, about $0.84 per AA task at max, list $3 / $15. The vendor README’s 2.8T / 104B is a different citation of the same scale; do not average the two. The AA tile is the published integer 60 plus its harness label. Do not average AA cost with DeepSWE $4.65/task.

07

Provenance

ClaimSourceAs of
DeepSWE 69% ±5% at maxDeepSWE official board2026-08-13
Official K3 cache-hit / miss / out $0.30 / $3.00 / $15.00 per 1M; context 1,048,576Kimi K3 pricing2026-08-19
OR catalog $3 / $15, cache read $0.30/M, hugging_face_id moonshotai/Kimi-K3OpenRouter catalog2026-08-19
HF safetensors.total 2,779,931,837,184; usedStorage 1,561,018,243,668; 96 shards; likes 10,848; downloads 2,289,863; license_name kimi-k3; dtypes F32 11,122,432 / BF16 57,179,884,544 / U8 2,722,740,830,208; 39 spacesHugging Face API moonshotai/Kimi-K32026-08-19
AA Index 60; 38.3 tok/s; 2800B / 104B; $0.84/AA task; $3 / $15Artificial Analysis Kimi K32026-08-19
vals Index 57.81% ±1.06; $6.468/test; 69 min 34 svals.ai Kimi K32026-08-19
Vendor README 2.8T / 104B activated; config KimiK3ForConditionalGeneration, hidden 7168HF README + config.json Kimi-K32026-08-19
Hosted id kimi-k3GET /v1/chat/hosted/models/public2026-08-19
1.32T week tokens (−9% WoW)OpenRouter rankings This Week2026-08-19
OpenRouter id moonshotai/kimi-k3, context 1,048,576OpenRouter /api/v1/models2026-08-19
HF downloads 2,289,863Hugging Face API moonshotai/Kimi-K32026-08-19
96 safetensor shards; shard bytes 1,560,936,091,448; tree files 118Hugging Face tree API moonshotai/Kimi-K32026-08-19
Spaces API returned 39; official collection Kimi K3 (2 items, 111 upvotes)Hugging Face Spaces / collections API2026-08-19
08

Compare, FAQ, and Get Plus

Same official DeepSWE harness (mini-swe-agent, public 113 tasks), fetched 2026-08-19. Different effort labels are the lab's own setting on that board, shown here rather than normalized.

Continue through Moonshot's model family with Kimi K2.7 Code.

FAQ

Why is the license “other”?

The Hub license field is other and license_name is kimi-k3. Read the official LICENSE in the repo. This page will not relabel it MIT.

Is 2.8T the Hub tensor count?

No. 2.8T / 104B is vendor README copy. Hub safetensors.total is 2,779,931,837,184. We cite both and do not average them.

Is thinking optional?

Official quickstart: thinking stays on. reasoning_effort is low / high / max, default max.

Why does the lab blog disagree with DeepSWE or Scale?

Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.

OpenAI-compatible call

Verified against the live public allowlist on 2026-08-19. Base URL is https://continuumcode.ai/v1. Keys are cont_sk_ from Settings, Account, Inference API. Personal keys need Plus or above.

curl https://continuumcode.ai/v1/chat/completions \
  -H "Authorization: Bearer $CONTINUUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"kimi-k3","messages":[{"role":"user","content":"Review this diff."}]}'

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.