- bench
- DeepSWE
- version
- v1.1
- split
- public 113 tasks
- harness
- mini-swe-agent
- effort
- high
- n
- 113
- metric
- Pass@1
- value
- 65%
- ci
- ±2%
- $/task
- $2.18
- tokens
- 107k out
- steps
- 125
Closed. Not listed as a Continuum hosted id. Hosted Flash is gemini-3.5-flash.
Use 3.7 Flash when you want the newest Google Flash and you accept first-party or OpenRouter billing. DeepSWE official: 65% ±2% at high, $2.18 per task, 107k out, 125 steps.
vals.ai Gemini 3.7 Flash card is cited as the 59.31% hero chip. We do not invent CI, $/test, latency, or named-bench percents that were not copied from that page into this card’s prior fetch notes.
Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.
Copied from AA or vals HTML we opened. 0.0% placeholder bars are omitted. Do not average these into the hero tiles, and do not invent CI, $/task, or steps for Artificial Analysis.
| Bench | Printed | Note | Source | URL | As of |
|---|---|---|---|---|---|
| AA list pair | $0.75 / $3.75 per 1M | Printed AA table. OpenRouter catalog is $0.375 / $1.875. Do not collapse the two, and do not paste Pro USD onto Flash. | Artificial Analysis Gemini 3.7 Flash | source | 2026-08-19 |
| AA throughput | 372.4 tok/s | Printed Speed row. Not a coding score. | Artificial Analysis Gemini 3.7 Flash | source | 2026-08-19 |
| AA task cost (their table) | about $0.40 per AA task | Printed AA cost footnote. Not DeepSWE $2.18/task. | Artificial Analysis Gemini 3.7 Flash | source | 2026-08-19 |
| OpenRouter list pair | $0.375 / $1.875 per 1M | Catalog row plus search $0.014 and cache fields cited in closedSurface. | OpenRouter catalog | source | 2026-08-19 |
Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.
Google does not publish Gemini Flash weights. No HF button.
Lab card and OpenRouter listing only. No Hugging Face Files button, because there is no official repo to point at.
Fetched OpenRouter catalog on 2026-08-19. hugging_face_id was empty.
Not listed as a Continuum hosted id. OpenRouter google/gemini-3.7-flash and first-party Gemini.
We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.
DeepSWE official: $2.18 per task at high (107k out, 125 steps). That is the hero economics number.
OpenRouter catalog 19 Aug 2026: $0.375 / $1.875 per million, web search $0.014, cache read $0.0375, cache write ≈ $0.02083 /M. Google lists $0.75 / $3.75 per million first party on its own pricing page, rising to $1.50 / $7.50 on 1 January 2027 (checked 20 Aug 2026); Artificial Analysis reprints the same figure. The OpenRouter row above is the aggregator rate.
Artificial Analysis footnote, 19 Aug 2026: $0.75 / $3.75 on their table, about $0.40 per AA task, 372.4 tok/s. The AA tile is Intelligence Index 56. AA did not publish a DeepSWE-style CI or step count we will copy onto the chip. Do not average AA cost with DeepSWE $2.18/task.
| Claim | Source | As of |
|---|---|---|
| DeepSWE 65% ±2% at high | DeepSWE official board | 2026-08-13 |
| OR catalog $0.375 / $1.875, search $0.014, cache read $0.0375, cache write ~$0.02083 | OpenRouter catalog | 2026-08-19 |
| Google pricing / 3.7 Flash pages opened; no extractable per-M table in saved server HTML/md | Gemini API pricing | 2026-08-19 |
| Not a Continuum hosted id; 3.5-flash is | GET /v1/chat/hosted/models/public | 2026-08-19 |
| OpenRouter id google/gemini-3.7-flash, context 1,048,576 | OpenRouter /api/v1/models | 2026-08-19 |
Same official DeepSWE harness (mini-swe-agent, public 113 tasks), fetched 2026-08-19. Different effort labels are the lab's own setting on that board, shown here rather than normalized.
Continue through Google's model family with Gemini 3.6 Flash.
No. The live public allowlist on 19 Aug 2026 names gemini-3.5-flash. 3.7 is first-party / OpenRouter only on that date.
Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.
No Continuum hosted id is verified for this model, so this page does not invent a curl target. Use the first-party API or the OpenRouter slug google/gemini-3.7-flash.