Models/Coding leaderboard

Best coding models:
the independent leaderboard

Three independent boards, stacked and never blended. 19 cards carry a DeepSWE official row, 22 carry a vals.ai hero index, 123 carry an Artificial Analysis integer. Every row is a chip that already exists on a model card, with its source URL and its fetch date.

Rows whose confidence interval overlaps another row are marked. On 113 tasks most of the top of the DeepSWE board is one band, not an ordering. Generated from the cards on 2026-08-19.

01

DeepSWE official, ranked by Pass@1

DeepSWE v1.1 Best, mini-swe-agent harness, public 113-task split, board updated 2026-08-13. Effort is the lab's own setting on that board and is printed rather than normalized. Cost, output tokens, and steps come from the same row as the pass rate, which is why this is the board we lead with: a percentage with no cost attached is not a purchasing decision.

19 rows, mini-swe-agent harness, public 113 tasksboard 2026-08-13
#ModelVendorPass@1$/taskOut tokensStepsEffortAs of
1 Claude Opus 5Not separated from #2, #3, #4, #5 anthropic 74%±4% $11.84 118k out 99 max 2026-08-13
2 GPT-5.6 SolNot separated from #1, #3, #4, #5 openai 73%±3% $8.39 60k out 61 max 2026-08-13
3 Claude Fable 5Not separated from #1, #2, #4, #5, #8, #9 anthropic 70%±4% $21.63 119k out 88 max 2026-08-13
4 Kimi K3Not separated from #1, #2, #3, #5, #8, #9 moonshotai 69%±5% $4.65 81k out 98 max 2026-08-13
5 GPT-5.5Tied on the printed value at #5. Not separated from #1, #2, #3, #4, #8, #9, #10 openai 67%±6% $7.23 46k out 82 xhigh 2026-08-19
5 GPT-5.6 LunaTied on the printed value at #5. Not separated from #1, #2, #3, #4, #8, #9 openai 67%±4% $0.61 73k out 102 max 2026-08-13
5 Grok 4.6Tied on the printed value at #5. Not separated from #3, #4, #8, #9 x-ai 67%±2% $5.50 71k out 87 xhigh 2026-08-13
8 Gemini 3.7 FlashNot separated from #3, #4, #5, #9 google 65%±2% $2.18 107k out 125 high 2026-08-13
9 DeepSeek V4 ProNot separated from #3, #4, #5, #8, #10, #11, #12, #13 deepseek 63%±6% $0.24 106k out 155 max 2026-08-13
10 Claude Opus 4.8Not separated from #5, #9, #11, #12, #13 anthropic 59%±2% $13.22 135k out 120 max 2026-08-19
11 Qwen3.8 MaxNot separated from #9, #10, #12, #13, #15 qwen 57%±3% $3.73 95k out 111 xhigh 2026-08-13
12 Muse Spark 1.2Not separated from #9, #10, #11, #13, #15 meta 55%±2% $3.70 99k out 101 xhigh 2026-08-13
13 DeepSeek V4 FlashTied on the printed value at #13. Not separated from #9, #10, #11, #12, #15, #16 deepseek 53%±4% $0.10 108k out 153 max 2026-08-13
13 Muse Spark 1.1Tied on the printed value at #13. Not separated from #11, #12, #15, #16 meta 53%±3% $2.36 74k out 96 xhigh 2026-08-19
15 GPT-5.4Not separated from #11, #12, #13, #16 openai 52%±2% $5.65 71k out 70 xhigh 2026-08-19
16 Gemini 3.6 FlashNot separated from #13, #15, #17 google 47%±4% not published not published not published high 2026-08-13
17 GLM 5.2Not separated from #16 z-ai 44%±2% $3.92 not published not published not published 2026-08-13
18 Gemini 3.5 FlashNot separated from #19 google 36%±4% $3.45 76k out 105 high 2026-08-19
19 Claude Sonnet 4.6Not separated from #18 anthropic 30%±4% $5.52 76k out 134 high 2026-08-19

19 rows: every card in this cluster carrying a DeepSWE official chip, and no others. Rails draw the printed pass rate on a fixed 100-point track, so bar length is the score itself and not a distance from the leader. A blank cell means the board did not print that field for that row, not that the value is zero. Board source.

02

vals.ai hero index, a different board

This is not a continuation of the table above. The vals.ai hero index is a separate methodology over a separate task mix, so a model can place well here and poorly on DeepSWE without either number being wrong. Ranks in this table are only comparable to other ranks in this table.

22 rows, vals.ai hero index, a separate methodologychecked 2026-08-19
#ModelVendorVals IndexAs of
1 GLM 5.3 z-ai 71.48% 2026-08-19
2 Claude Opus 5 anthropic 67.21% 2026-08-19
3 Claude Fable 5 anthropic 66.04% 2026-08-19
4 GPT-5.6 Sol openai 63.71% 2026-08-19
5 Gemini 3.5 Flash google 62.05% 2026-08-19
6 GPT-5.6 Luna openai 59.88% 2026-08-19
7 Claude Sonnet 5 anthropic 59.61% 2026-08-19
8 Gemini 3.7 Flash google 59.31% 2026-08-19
9 Kimi K3 moonshotai 57.81% 2026-08-19
10 GPT-5.6 Terra openai 56.53% 2026-08-19
11 Gemini 3.6 Flash google 55.35% 2026-08-19
12 DeepSeek V4 Flash deepseek 53.57% 2026-08-19
13 DeepSeek V4 Pro deepseek 52.37% 2026-08-19
14 Qwen3.8 Max qwen 51.84% 2026-08-19
15 Command A cohere 43.41% 2026-08-19
16 MiniMax M3 minimax 42.72% 2026-08-19
17 Gemini 3.1 Pro Preview google 41.90% 2026-08-19
18 MiMo-V2.5-Pro xiaomi 40.97% 2026-08-19
19 MiMo-V2.5 xiaomi 39.91% 2026-08-19
20 Inkling thinkingmachines 34.1% 2026-08-19
21 Inkling Small thinkingmachines 31.86% 2026-08-19
22 Nemotron 3 Ultra nvidia 27.39% 2026-08-19

22 rows. vals.ai prints an index percent per card; where it also printed an interval, that interval lives on the model card, not here. Board source.

03

Artificial Analysis Intelligence Index, a third board

A third board again, and the least coding-specific of the three: the Intelligence Index is a published integer over nine evaluations, most of which are not software engineering. It is here because it is independent and broadly available, not because it answers the coding question. Artificial Analysis publishes no interval, so equal integers are genuinely not separated and are shown sharing a rank.

123 rows, published integer over nine evaluationschecked 2026-08-19
#ModelVendorIndexAs of
1 Claude Opus 5 anthropic 63 2026-08-19
2 Claude Fable 5 anthropic 62 2026-08-19
3 GPT-5.6 SolTied on the printed integer openai 61 2026-08-19
3 Grok 4.6Tied on the printed integer x-ai 61 2026-08-19
5 GLM 5.3Tied on the printed integer z-ai 60 2026-08-19
5 Kimi K3Tied on the printed integer moonshotai 60 2026-08-19
7 Qwen3.8 2.4T A95BTied on the printed integer qwen 58 2026-08-19
7 Qwen3.8 MaxTied on the printed integer qwen 58 2026-08-19
9 Claude Opus 4.8Tied on the printed integer anthropic 57 2026-08-19
9 GPT-5.6 TerraTied on the printed integer openai 57 2026-08-19
9 Muse Spark 1.2Tied on the printed integer meta 57 2026-08-19
12 Gemini 3.7 FlashTied on the printed integer google 56 2026-08-19
12 GPT-5.5Tied on the printed integer openai 56 2026-08-19
12 Grok 4.5Tied on the printed integer x-ai 56 2026-08-19
15 Claude Opus 4.7Tied on the printed integer anthropic 55 2026-08-19
15 Claude Sonnet 5Tied on the printed integer anthropic 55 2026-08-19
17 DeepSeek V4 ProTied on the printed integer deepseek 53 2026-08-19
17 GLM 5.2Tied on the printed integer z-ai 53 2026-08-19
17 GPT-5.4Tied on the printed integer openai 53 2026-08-19
17 Muse Spark 1.1Tied on the printed integer meta 53 2026-08-19
21 DeepSeek V4 FlashTied on the printed integer deepseek 52 2026-08-19
21 Gemini 3.5 FlashTied on the printed integer google 52 2026-08-19
21 Gemini 3.6 FlashTied on the printed integer google 52 2026-08-19
21 GPT-5.6 LunaTied on the printed integer openai 52 2026-08-19
21 Qwen3.8 27BTied on the printed integer qwen 52 2026-08-19
26 Qwen3.7 Max qwen 47 2026-08-19
27 GPT-5.3-Codex openai 46 2026-08-19
28 Kimi K2.6Tied on the printed integer moonshotai 45 2026-08-19
28 MiniMax M3Tied on the printed integer minimax 45 2026-08-19
30 GPT-5.2Tied on the printed integer openai 43 2026-08-19
30 Kimi K2.7 CodeTied on the printed integer moonshotai 43 2026-08-19
30 MiMo-V2.5-ProTied on the printed integer xiaomi 43 2026-08-19
33 Hy3Tied on the printed integer tencent 42 2026-08-19
33 InklingTied on the printed integer thinkingmachines 42 2026-08-19
33 Solar Pro 4Tied on the printed integer upstage 42 2026-08-19
36 GLM 5Tied on the printed integer z-ai 41 2026-08-19
36 GLM 5.1Tied on the printed integer z-ai 41 2026-08-19
36 GPT-5.2-CodexTied on the printed integer openai 41 2026-08-19
36 GPT-5.4 MiniTied on the printed integer openai 41 2026-08-19
36 Inkling SmallTied on the printed integer thinkingmachines 41 2026-08-19
41 GPT-5.4 NanoTied on the printed integer openai 40 2026-08-19
41 Qwen3.6 PlusTied on the printed integer qwen 40 2026-08-19
43 Claude Opus 4.6Tied on the printed integer anthropic 39 2026-08-19
43 GLM 5 TurboTied on the printed integer z-ai 39 2026-08-19
43 MiniMax M2.7Tied on the printed integer minimax 39 2026-08-19
43 Qwen3.7 PlusTied on the printed integer qwen 39 2026-08-19
47 Grok 4.3Tied on the printed integer x-ai 38 2026-08-19
47 Ling 3.0 FlashTied on the printed integer inclusionai 38 2026-08-19
47 MiMo-V2.5Tied on the printed integer xiaomi 38 2026-08-19
47 Nemotron 3 UltraTied on the printed integer nvidia 38 2026-08-19
47 Qwen3.6 27BTied on the printed integer qwen 38 2026-08-19
52 Claude Sonnet 4.6Tied on the printed integer anthropic 37 2026-08-19
52 Gemini 3.5 Flash LiteTied on the printed integer google 37 2026-08-19
52 GPT-5.1Tied on the printed integer openai 37 2026-08-19
55 Claude Opus 4.5Tied on the printed integer anthropic 36 2026-08-19
55 GPT-5.1-CodexTied on the printed integer openai 36 2026-08-19
55 Kimi K2.5Tied on the printed integer moonshotai 36 2026-08-19
58 GLM 5V TurboTied on the printed integer z-ai 35 2026-08-19
58 GPT-5Tied on the printed integer openai 35 2026-08-19
58 Qwen3.5-27BTied on the printed integer qwen 35 2026-08-19
61 GLM 4.7Tied on the printed integer z-ai 34 2026-08-19
61 KAT-Coder-Pro V2Tied on the printed integer kwaipilot 34 2026-08-19
61 MiniMax M2.5Tied on the printed integer minimax 34 2026-08-19
61 Qwen3.5 397B A17BTied on the printed integer qwen 34 2026-08-19
61 Tencent HY3 PreviewTied on the printed integer tencent 34 2026-08-19
66 Kimi K2 ThinkingTied on the printed integer moonshotai 33 2026-08-19
66 o3 ProTied on the printed integer openai 33 2026-08-19
66 Qwen3.5-122B-A10BTied on the printed integer qwen 33 2026-08-19
69 MiniMax M2.1Tied on the printed integer minimax 32 2026-08-19
69 Qwen3 Max ThinkingTied on the printed integer qwen 32 2026-08-19
69 Qwen3.6 35B A3BTied on the printed integer qwen 32 2026-08-19
69 Ring-2.6-1TTied on the printed integer inclusionai 32 2026-08-19
73 GPT-5.1-Codex-MiniTied on the printed integer openai 31 2026-08-19
73 o3Tied on the printed integer openai 31 2026-08-19
73 Step 3.7 FlashTied on the printed integer stepfun 31 2026-08-19
76 Mistral Medium 3.5Tied on the printed integer mistralai 30 2026-08-19
76 Qwen3.5-35B-A3BTied on the printed integer qwen 30 2026-08-19
78 MiniMax M2 minimax 29 2026-08-19
79 Ling 2.6 1TTied on the printed integer inclusionai 27 2026-08-19
79 Step 3.5 FlashTied on the printed integer stepfun 27 2026-08-19
81 Gemini 2.5 ProTied on the printed integer google 26 2026-08-19
81 GPT-5 MiniTied on the printed integer openai 26 2026-08-19
81 Nemotron 3 Super (free)Tied on the printed integer nvidia 26 2026-08-19
81 o4 MiniTied on the printed integer openai 26 2026-08-19
85 DeepSeek V3.2 deepseek 25 2026-08-19
86 gpt-oss-120bTied on the printed integer openai 24 2026-08-19
86 Kimi K2 0905Tied on the printed integer moonshotai 24 2026-08-19
86 Nemotron 3.5 LightningTied on the printed integer nvidia 24 2026-08-19
86 Qwen3 MaxTied on the printed integer qwen 24 2026-08-19
90 GLM 4.6Tied on the printed integer z-ai 23 2026-08-19
90 GLM 4.7 FlashTied on the printed integer z-ai 23 2026-08-19
92 DeepSeek V3.1 TerminusTied on the printed integer deepseek 22 2026-08-19
92 Qwen3.5-9BTied on the printed integer qwen 22 2026-08-19
94 Qwen3 Coder Next qwen 21 2026-08-19
95 GPT-4.1Tied on the printed integer openai 20 2026-08-19
95 GPT-5 NanoTied on the printed integer openai 20 2026-08-19
95 Kimi K2 0711Tied on the printed integer moonshotai 20 2026-08-19
95 North Mini CodeTied on the printed integer cohere 20 2026-08-19
95 R1 0528Tied on the printed integer deepseek 20 2026-08-19
100 Sonar Reasoning Pro perplexity 18 2026-08-19
101 Mistral Large 3 mistralai 16 2026-08-19
102 GPT-4.1 MiniTied on the printed integer openai 15 2026-08-19
102 gpt-oss-20bTied on the printed integer openai 15 2026-08-19
102 Mistral Medium 3.1Tied on the printed integer mistralai 15 2026-08-19
105 Gemini 2.5 FlashTied on the printed integer google 14 2026-08-19
105 Ling-2.6-flashTied on the printed integer inclusionai 14 2026-08-19
105 Llama 4 MaverickTied on the printed integer meta 14 2026-08-19
105 Solar Pro 3Tied on the printed integer upstage 14 2026-08-19
109 Amazon Nova Premier amazon 13 2026-08-19
110 Mistral Medium 3 mistralai 12 2026-08-19
111 GLM 4.6V z-ai 11 2026-08-19
112 GPT-4.1 NanoTied on the printed integer openai 10 2026-08-19
112 Llama 4 ScoutTied on the printed integer meta 10 2026-08-19
112 R1 Distill Llama 70BTied on the printed integer deepseek 10 2026-08-19
115 Hermes 4Tied on the printed integer nousresearch 9 2026-08-19
115 SonarTied on the printed integer perplexity 9 2026-08-19
117 Command ATied on the printed integer cohere 7 2026-08-19
117 Gemini 2.5 Flash LiteTied on the printed integer google 7 2026-08-19
117 GLM 4.5VTied on the printed integer z-ai 7 2026-08-19
117 Nemotron Nano 9B V2 (free)Tied on the printed integer nvidia 7 2026-08-19
121 Saba mistralai 6 2026-08-19
122 Hermes 3 70B Instruct nousresearch 5 2026-08-19
123 Nemotron Nano 12B 2 VL (free) nvidia 4 2026-08-19

123 rows. The integer is the whole chip: per-evaluation percents printed on an AA model page sit in that card's named-benches table and are never averaged into anything. Board source.

04

How this leaderboard is built

Every row on this page comes from a chip that already exists on a Continuum model card, and every chip carries a source URL and a fetch date. Nothing here is computed by us, averaged by us, or estimated. If a board did not publish a number for a model, that model has no row.

What DeepSWE measures. The official DeepSWE board runs mini-swe-agent, a deliberately minimal bash-only agent, against 113 public software-engineering tasks. The reward is binary: the patch either makes the repository's tests pass or it does not. There is no partial credit, no human grader, and no model acting as judge. Because the harness is fixed and published, the board can also report what a run cost, which is where the dollars per task, output tokens, and step count in the first table come from. Effort is the lab's own setting on that board. A row at max and a row at high are not the same experiment, so we print the label instead of hiding it.

Why the intervals matter more than the order. 113 tasks is a small set. The board publishes an interval around each pass rate, and on this data most of the top rows overlap. When two intervals overlap, the board cannot tell you which model is better, only that both sit in the same band. Rather than sell a clean one, two, three, we mark every row that overlaps another and name the ranks it cannot be separated from. Read the bands, not the ordering.

Why we do not blend boards. The three tables above are three different experiments. DeepSWE is a coding-agent pass rate on a fixed harness. The Artificial Analysis Intelligence Index is a published integer over nine evaluations. The vals.ai hero index is a third methodology on a third task mix. Averaging them produces a number nobody measured and nobody can reproduce. We stack the boards. We do not merge them, and we do not offer a combined score.

Why usage rankings are not on this page. OpenRouter This Week ranks models by tokens routed. That is popularity: it tracks price, free tiers, the default model in somebody else's tool, and which model a large customer wired up last month. A cheap model serving a high-volume classification job outranks every frontier coder on it. We publish that board on the models index because it is useful market data, and we keep it off this page because it is not evidence of quality.

Why vendor self-reports are excluded. A lab that picks its own harness, its own effort setting, its own subset, and its own retry policy is not running the same test as the board. Vendor figures appear on individual cards, labeled vendor-reported, and they never enter a ranking here. The same rule removes boards we could not fetch cleanly: Scale SWE-bench Pro and Terminal-Bench 2.1 are omitted cluster-wide rather than borrowed from a scraper.

05

FAQ

What is the best coding model right now?

On the DeepSWE official board we fetched, the highest printed row is Claude Opus 5 at 74% ±4% Pass@1 at max, board date 2026-08-13. It is not a separated win. Its interval overlaps 5 other rows (GPT-5.6 Sol, Claude Fable 5, Kimi K3, GPT-5.5, GPT-5.6 Luna), so on 113 tasks the board cannot tell you which of them is better. The honest way to pick inside that band is cost: it runs from $0.61 per task on GPT-5.6 Luna to $21.63 on Claude Fable 5, a gap the pass rate does not justify. Anyone naming a single best coding model without an interval is selling you an ordering the data does not support.

Why is SWE-bench Verified not used?

SWE-bench Verified is historical here and is never a hero chip. The scores in circulation for it were produced by different scaffolds, different retry budgets, and different dates, so putting them in one column would imply a comparison nobody ran. DeepSWE publishes one fixed harness (mini-swe-agent), one split, and the cost of each run, which is what makes a column meaningful. Where a card cites SWE-bench Verified at all, it sits in that card's named-benches table as a printed figure with its source, not as an input to any ranking on this page.

Are these scores independent?

Yes, and that is the entry condition. Every row comes from a board run by someone other than the lab being scored: DeepSWE at deepswe.datacurve.ai, Artificial Analysis, and vals.ai. Vendor-reported figures are excluded from all three tables. The cluster currently carries one vendor-reported eval row; it is labeled vendor-reported on its own card and it does not appear here.

How often does this update?

This page is generated from the model cards, so it moves when a chip moves and it cannot drift away from them. Current fetch dates: the DeepSWE board at 2026-08-13, Artificial Analysis and vals.ai at 2026-08-19. A model with no row is absent rather than estimated, so the row count changing is itself the signal that a board published something new.

After the board

Run the band, not the ranking.

Continuum hosts several of the models above under one id and one key, so you can put two rows from the same band on the same task and read your own numbers.

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.