- bench
- DeepSWE
- version
- v1.1
- split
- public 113 tasks
- harness
- mini-swe-agent
- effort
- high
- n
- 113
- metric
- Pass@1
- value
- 47%
- ci
- ±4%
Closed. OpenRouter catalog 19 Aug 2026: $0.75 / $3.75, cache read $0.075 /M, cache write ≈$0.0417 /M, web search $0.014/request, 1,048,576 / 65,536, created unix 1784646733.
Newest Flash on this hub is 3.7 (830B, marked new, $0.25 / $1.25 on OR). Continuum’s live hosted Flash id is still gemini-3.5-flash.
Use 3.6 Flash when you specifically want this checkpoint. Official DeepSWE mini-swe-agent at high: 47% ±4: the same 13 Aug family note already published on the 3.7 Flash hub line. AA 52. vals 55.35% ±1.09 is a different harness.
vals.ai Accuracy Rankings hero, 19 Aug 2026: 55.35% ±1.09. Different harness from DeepSWE 47% ±4.
Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.
Copied from AA or vals HTML we opened. 0.0% placeholder bars are omitted. Do not average these into the hero tiles, and do not invent CI, $/task, or steps for Artificial Analysis.
Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.
Google does not publish 3.6 Flash weights. No HF Files button.
Lab card and OpenRouter listing only. No Hugging Face Files button, because there is no official repo to point at.
Fetched OpenRouter catalog on 2026-08-19. hugging_face_id was empty.
Not listed as a Continuum hosted id. Pair a local Google key or OpenRouter google/gemini-3.6-flash.
We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.
No DeepSWE $/task printed on the 13 Aug family note we cite for the 47% ±4 row.
OpenRouter + AA, 19 Aug 2026: $0.75 / $3.75 per million. Google lists 3.7 Flash on the same $0.75 / $3.75 through 31 December 2026, so there is no separate 3.7 figure to confuse this with.
Artificial Analysis Intelligence Index 52, fetched 19 Aug 2026. Same page: 203 tok/s; $412.03 Index eval. The tile is the published integer 52.
| Claim | Source | As of |
|---|---|---|
| DeepSWE 47% ±4% at high | DeepSWE official board | 2026-08-13 |
| OR catalog $0.75 / $3.75, cache $0.075 / ~$0.0417 write, search $0.014, 1,048,576 / 65,536, created 1784646733 | OpenRouter catalog | 2026-08-19 |
| DeepSWE 47% ±4 at high | DeepSWE official board (13 Aug family note) | 2026-08-13 |
| AA Index 52; $0.75 / $3.75; 203 tok/s; $412.03 Index eval | Artificial Analysis Gemini 3.6 Flash | 2026-08-19 |
| vals Accuracy 55.35% ±1.09 | vals.ai Accuracy Rankings 2026-08-19 | 2026-08-19 |
| 2.75T week tokens (+12% WoW) | OpenRouter rankings This Week | 2026-08-19 |
| OpenRouter id google/gemini-3.6-flash, context 1,048,576 | OpenRouter /api/v1/models | 2026-08-19 |
Same official DeepSWE harness (mini-swe-agent, public 113 tasks), fetched 2026-08-19. Different effort labels are the lab's own setting on that board, shown here rather than normalized.
Continue through Google's model family with Gemini 3.1 Pro Preview.
No. OR 19 Aug prices 3.6 at $0.75 / $3.75 and 3.7 at $0.25 / $1.25.
47% ±4 is 3.6 Flash at high. 49% ±3 is 3.7 Flash. Same official board, different SKUs.
Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.
No Continuum hosted id is verified for this model, so this page does not invent a curl target. Use the first-party API or the OpenRouter slug google/gemini-3.6-flash.