Models·25 labs

The coding-agent
model dossier.

A card is one canonical model identity: aliases, whether Continuum hosts it, and any coding-agent score we could fetch with the harness still visible. Vendor copy is labeled vendor. Independent boards are labeled independent. We omit a chip rather than invent a percent.

Ranked view Best coding models: the independent leaderboard Every card carrying an independent coding score on one page: 19 DeepSWE official rows ranked by Pass@1 with cost per task and steps, then vals.ai and Artificial Analysis as separate boards that are never blended into it. Overlapping confidence intervals are marked rather than hidden behind a clean ranking. Open the leaderboard
01

What a card is

Each card is one identity, not a SKU dump. This cluster publishes 223 cards across 25 labs (at least two per vendor). OpenRouter This Week tokens and official Hugging Face downloads are the two sort keys on this index. Preview dates that did not earn a card stay on the vendor hub as lineage. Fable 5 is a section in Claude Code models; it also has a full card here because buyers search the name.

Vendor means the lab that ships the weights or the API. Independent means a board we opened: DeepSWE official mini-swe-agent (v1.1 Best, 113 tasks, board updated 13 Aug 2026), Artificial Analysis Intelligence Index integers, and vals.ai hero-index percents where we opened a card. Scale SWE-bench Pro and Terminal-Bench 2.1 stay omitted cluster-wide.

Open-weight cards link the official Hugging Face repo (Files and Use this model). Closed-weight cards have no fake download. Continuum hosted ids come from the live public allowlist on 19 Aug 2026, not from a guess.

02

This Week

Source of record: OpenRouter This Week, usage through 18 Aug 2026, fetched 19 Aug 2026. Tokens are prompt plus completion. The percent on a leaderboard row is week-over-week change, not share of the week. Market share on this page is text request share by author.

OpenRouter This Week tokens are prompt plus completion. The percent next to a row is week-over-week change, not share of the week. Market share on this page is text request share by author, a different number.

20 rows, tokens routed, This Weekfetched 2026-08-19
#ModelAuthorTokensWoW
1 DeepSeek V4 Flash 0731 deepseek 11.3T +12%
2 Hy3 tencent 9.83T +4%
3 GPT-5.6 Luna openai 5.8T +19%
4 MiMo-V2.5 xiaomi 5.46T +9%
5 DeepSeek V4 Flash 0423 deepseek 4.78T −13%
6 GLM 5.2 z-ai 4.34T +18%
7 Gemini 3.6 Flash google 2.75T +12%
8 Nemotron 3 Ultra (free) nvidia 2.69T +20%
9 Claude Opus 5 anthropic 2.68T +89%
10 DeepSeek V4 Pro 0423 deepseek 2.57T +1%
11 MiniMax M3 minimax 1.67T +2%
12 Laguna S 2.1 (free) poolside 1.67T −5%
13 Kimi K3 moonshotai 1.32T −9%
14 Claude Sonnet 5 anthropic 1.08T +5%
15 DeepSeek V4 Pro 0813 deepseek 946B new
16 GPT-5.6 Terra openai 943B +12%
17 Step 3.7 Flash stepfun 857B −30%
18 Nemotron 3.5 Lightning (free) nvidia 836B +999%
19 Gemini 3.7 Flash google 830B new
20 Gemini 3 Flash Preview google 821B −8%

Rails on this board are relative: token volume has no natural ceiling, so each fill is that row's tokens against the largest row here. Read them against each other, not as a percentage.

Second sort key: official Hugging Face downloads on the text-generation and image-text-to-text repos we carded. Hub API 2026-08-19. Popularity, not quality. Embedding, MiniLM, BERT, BGE, CLIP, and TTS boards are omitted.

Top 20 carded repos by Hub downloadshub api 2026-08-19
#ModelAuthorDownloadsHub
1 Qwen3-0.6B qwen 28,202,544 Qwen/Qwen3-0.6B
2 Qwen3 8B qwen 15,796,910 Qwen/Qwen3-8B
3 Qwen3.5-9B qwen 13,604,767 Qwen/Qwen3.5-9B
4 Qwen2.5-7B-Instruct qwen 12,229,058 Qwen/Qwen2.5-7B-Instruct
5 Gemma 4 26B A4B google 9,533,369 google/gemma-4-26B-A4B-it
6 Gemma 4 31B (free) google 9,311,525 google/gemma-4-31B-it
7 Llama 3.2 1B Instruct meta 8,694,182 meta-llama/Llama-3.2-1B-Instruct
8 gpt-oss-20b openai 7,682,588 openai/gpt-oss-20b
9 Llama 3.1 8B Instruct meta 7,199,331 meta-llama/Meta-Llama-3.1-8B-Instruct
10 R1 deepseek 6,911,569 deepseek-ai/DeepSeek-R1
11 Qwen3.6 27B qwen 6,745,154 Qwen/Qwen3.6-27B
12 Qwen3.6 35B A3B qwen 5,663,515 Qwen/Qwen3.6-35B-A3B
13 Qwen3 VL 8B Instruct qwen 5,280,026 Qwen/Qwen3-VL-8B-Instruct
14 Gemma 3 1B IT google 4,983,962 google/gemma-3-1b-it
15 gpt-oss-120b openai 4,657,776 openai/gpt-oss-120b
16 Qwen3.5-27B qwen 2,837,939 Qwen/Qwen3.5-27B
17 GLM 5.2 z-ai 2,748,563 zai-org/GLM-5.2
18 Qwen3.5-35B-A3B qwen 2,427,827 Qwen/Qwen3.5-35B-A3B
19 DeepSeek V4 Flash deepseek 2,330,940 deepseek-ai/DeepSeek-V4-Flash-0731
20 Kimi K3 moonshotai 2,289,863 moonshotai/Kimi-K3

/models sorted Top Weekly extends past the 20-row LLM Leaderboard and uses slightly different totals because it counts all endpoints. Sol is #18 there at 973B while Lightning is #18 on the leaderboard. We cite both lists and do not average them.

GPT-5.6 Sol973B#18 on /models Top Weekly
Solar Pro 4497B#29
Grok 4.6490B#30
Claude Fable 5310B#38
Qwen3.8 Max297B#39
North Mini Code238B#42
Grok 4.5236B#43
03

Request share

Market share is text request count by author, not tokens. meta-llama is folded onto the Meta hub. The +12% on Flash 0731 is week-over-week change, not this 26.3% request share.

10 authors, text requests, This Weekfetched 2026-08-19
AuthorRequestsShare
deepseek 418M 26.3%
google 370M 23.2%
openai 288M 18.1%
qwen 90.8M 5.7%
anthropic 59.5M 3.7%
xiaomi 53.2M 3.3%
tencent 52.8M 3.3%
meta-llama 44.1M 2.8%
mistralai 41.2M 2.6%
Others 174M 11.0%
04

25 vendors

Hub order is locked: unique authors, token-first from the This Week board, then request share, then named labs. Moonshot and StepFun stay because they are on the week board. xAI, Meta, and Qwen stay as named plus request-share labs. Week-token cells name the SKU we actually counted. We do not invent a figure for a lab the boards did not cite.

01 DeepSeekFlash 0731 leads the week board DeepSeek V4 Flash 11.3TDeepSeek V4 Flash 0731 02 TencentHy3, second on the week board Hy3 9.83THy3 03 OpenAILuna is the week SKU; Sol is the sibling card GPT-5.6 Luna 5.8TGPT-5.6 Luna 04 XiaomiMiMo-V2.5, #4 on the week board MiMo-V2.5 5.46TMiMo-V2.5 05 Z.aiGLM 5.3 is current; 5.2 was week volume GLM 5.3 4.34TGLM 5.2 06 Google3.7 Flash is newest; 3.6 led the week Gemini 3.7 Flash 2.75TGemini 3.6 Flash 07 NVIDIANemotron 3 Ultra, free lane Nemotron 3 Ultra 2.69TNemotron 3 Ultra (free) 08 AnthropicOpus 5 and Fable 5 both get full cards Claude Opus 5 2.68TClaude Opus 5 09 MiniMaxM3, #11 on the week board MiniMax M3 1.67TMiniMax M3 10 PoolsideLaguna S 2.1, #12 on the week board Laguna S 2.1 1.67TLaguna S 2.1 (free) 11 MoonshotKimi K3, #13 on the week board Kimi K3 1.32TKimi K3 12 StepFunStep 3.7 Flash, #17 on the week board Step 3.7 Flash 857BStep 3.7 Flash 13 UpstageSolar Pro 4, #29 on Top Weekly Solar Pro 4 497BSolar Pro 4 14 xAIGrok 4.6, #30 on Top Weekly Grok 4.6 490BGrok 4.6 15 QwenQwen3.8 Max; 5.7% request share Qwen3.8 Max 297BQwen3.8 Max 16 CohereCommand A; North Mini Code is week volume Command A 238BNorth Mini Code 17 MetaMuse Spark 1.2; Llama 4 request share lives here Muse Spark 1.2 2.8%request share 18 MistralLarge 3 2512; 2.6% request share Mistral Large 3 2.6%request share 19 ByteDance SeedSeed 2.1 Turbo; Code is the sibling Seed 2.1 Turbo No week tokens citedweek tokens unknown 20 AmazonNova 2 Lite is the only Nova 2 on OR Nova 2 Lite No week tokens citedweek tokens unknown 21 PerplexitySonar Pro Sonar Pro No week tokens citedweek tokens unknown 22 Nous ResearchHermes 4 405B Hermes 4 No week tokens citedweek tokens unknown 23 InclusionAILing 3.0 Flash Ling 3.0 Flash No week tokens citedweek tokens unknown 24 KwaipilotKAT Coder Pro v2.5 KAT Coder Pro v2.5 No week tokens citedweek tokens unknown 25 Thinking MachinesInkling Inkling No week tokens citedweek tokens unknown
05

How we score coding

Hero tiles are DeepSWE (official v1.1 Best, mini-swe-agent, 113 tasks, board updated 13 Aug 2026; effort labeled), Artificial Analysis Intelligence Index (published integer), and the vals.ai hero index percent where we opened a card. Missing benches are omitted, not invented. Every chip carries source URL and as-of.

Scale SWE-bench Pro is cited only from the public shared-harness board. We could not extract a published per-model row from labs.scale.com on 19 Aug 2026 (client-rendered). Terminal-Bench 2.1 is pinned to Artificial Analysis; their model pages did not expose a standalone Terminal-Bench 2.1 score in server HTML. Both stay omitted cluster-wide.

LiveCodeBench is a contest footnote if we ever have a sourced row. SWE-bench Verified is historical and never the hero. Artificial Analysis cost and speed, when fetched, sit under Economics as a footnote, not a substitute for the Intelligence Index integer.

After the honest block

Get Plus is the ask.

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.