- bench
- DeepSWE
- version
- v1.1
- split
- public 113 tasks
- harness
- mini-swe-agent
- effort
- high
- n
- 113
- metric
- Pass@1
- value
- 36%
- ci
- ±4%
- $/task
- $3.45
- tokens
- 76k out
- steps
- 105
gemini-3.5-flashContext window on the 2026-08-19 OpenRouter row: 1,048,576 tokens.
Max completion tokens on that row: 65,536.
Input / output on that row: $1.50 / $9.00 per 1M tokens.
Architecture fields: text+image+file+audio+video->text · Gemini.
No hugging_face_id on the 2026-08-19 OpenRouter row.
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed.
When you want the 2026-05-19 catalog SKU, not a later rename.
Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.
Copied from AA or vals HTML we opened. 0.0% placeholder bars are omitted. Do not average these into the hero tiles, and do not invent CI, $/task, or steps for Artificial Analysis.
| Bench | Printed | Note | Source | URL | As of |
|---|---|---|---|---|---|
| GDPval-AA v2 | 42.2% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
| τ³-Banking | 32.2% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
| Terminal-Bench v2.1 | 78.7% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
| SciCode | 53.1% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
| Humanity's Last Exam | 42.7% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
| GPQA Diamond | 92.2% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
| CritPt | 13.1% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
| AA-Omniscience Accuracy | 51.4% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
| AA-LCR | 81% | Printed on the AA Intelligence Evaluations grid for Gemini 3.5 Flash (high). Not a DeepSWE chip. Tile effort: high. | Artificial Analysis Gemini 3.5 Flash | source | 2026-08-19 |
gemini-3.5-flash. Use a different card on this hub when you need another lab SKU.Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.
Closed API SKU. The 2026-08-19 OpenRouter row lists pricing and context; there is no official Hugging Face weight dump on that row.
Lab card and OpenRouter listing only. No Hugging Face Files button, because there is no official repo to point at.
Fetched OpenRouter catalog on 2026-08-19. hugging_face_id was empty.
OpenRouter slug google/gemini-3.5-flash on the 2026-08-19 catalog.
First party: https://ai.google.dev/gemini-api/docs/models.
Continuum hosts gemini-3.5-flash on the live public allowlist fetched 2026-08-19.
gemini-3.5-flash
click to select
We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.
OpenRouter 2026-08-19: $1.50 in / $9.00 out per 1M tokens. Context 1,048,576 in / 65,536 out.
Internal reasoning on that row: $9.00 per 1M.
| Claim | Source | As of |
|---|---|---|
| DeepSWE 36% ±4% at high | DeepSWE official board | 2026-08-19 |
| OpenRouter id google/gemini-3.5-flash; context 1048576; created 1779193800; $1.50 / $9.00 per 1M | OpenRouter /api/v1/models | 2026-08-19 |
| DeepSWE 36% ±4% at high; $3.45/task | DeepSWE official mini-swe-agent board | 2026-08-19 |
| AA Intelligence Index 52; 9 printed Index benches | Artificial Analysis model page | 2026-08-19 |
| vals hero 62.05% | vals.ai model card | 2026-08-19 |
| OpenRouter id google/gemini-3.5-flash, context 1,048,576 | OpenRouter /api/v1/models | 2026-08-19 |
No other card in this cluster shares a published DeepSWE official row we can put next to this one. We will not compare on vendor-blog numbers.
Continue through Google's model family with Gemma 4 31B (free).
Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.
Verified against the live public allowlist on 2026-08-19. Base URL is https://continuumcode.ai/v1. Keys are cont_sk_ from Settings, Account, Inference API. Personal keys need Plus or above.
curl https://continuumcode.ai/v1/chat/completions \
-H "Authorization: Bearer $CONTINUUM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gemini-3.5-flash","messages":[{"role":"user","content":"Review this diff."}]}'