- bench
- DeepSWE
- version
- v1.1
- split
- public 113 tasks
- harness
- mini-swe-agent
- effort
- xhigh
- n
- 113
- metric
- Pass@1
- value
- 52%
- ci
- ±2%
- $/task
- $5.65
- tokens
- 71k out
- steps
- 70
Context window on the 2026-08-19 OpenRouter row: 1,050,000 tokens.
Max completion tokens on that row: 128,000.
Input / output on that row: $2.50 / $15.00 per 1M tokens.
Architecture fields: text+image+file->text · GPT.
No hugging_face_id on the 2026-08-19 OpenRouter row.
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system.
When you want the 2026-03-05 catalog SKU, not a later rename.
We did not open a vals.ai card for this identity. Chip omitted.
Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.
Copied from AA or vals HTML we opened. 0.0% placeholder bars are omitted. Do not average these into the hero tiles, and do not invent CI, $/task, or steps for Artificial Analysis.
| Bench | Printed | Note | Source | URL | As of |
|---|---|---|---|---|---|
| GDPval-AA v2 | 44.7% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
| τ³-Banking | 39.6% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
| Terminal-Bench v2.1 | 78.3% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
| SciCode | 56.6% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
| Humanity's Last Exam | 43.7% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
| GPQA Diamond | 92% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
| CritPt | 23.4% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
| AA-Omniscience Accuracy | 50.8% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
| AA-LCR | 77.7% | Printed on the AA Intelligence Evaluations grid for GPT-5.4 (xhigh). Not a DeepSWE chip. Tile effort: xhigh. | Artificial Analysis GPT-5.4 (xhigh) | source | 2026-08-19 |
Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.
Closed API SKU. The 2026-08-19 OpenRouter row lists pricing and context; there is no official Hugging Face weight dump on that row.
Lab card and OpenRouter listing only. No Hugging Face Files button, because there is no official repo to point at.
Fetched OpenRouter catalog on 2026-08-19. hugging_face_id was empty.
OpenRouter slug openai/gpt-5.4 on the 2026-08-19 catalog.
First party: https://platform.openai.com/docs/models.
Not on the Continuum host list fetched 2026-08-19.
We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.
OpenRouter 2026-08-19: $2.50 in / $15.00 out per 1M tokens. Context 1,050,000 in / 128,000 out.
No separate internal-reasoning price on that row.
| Claim | Source | As of |
|---|---|---|
| DeepSWE 52% ±2% at xhigh | DeepSWE official board | 2026-08-19 |
| OpenRouter id openai/gpt-5.4; context 1050000; created 1772734352; $2.50 / $15.00 per 1M | OpenRouter /api/v1/models | 2026-08-19 |
| DeepSWE 52% ±2% at xhigh; $5.65/task | DeepSWE official mini-swe-agent board | 2026-08-19 |
| AA Intelligence Index 53 (xhigh); 9 printed Index benches | Artificial Analysis model page | 2026-08-19 |
| OpenRouter id openai/gpt-5.4, context 1,050,000 | OpenRouter /api/v1/models | 2026-08-19 |
No other card in this cluster shares a published DeepSWE official row we can put next to this one. We will not compare on vendor-blog numbers.
Continue through OpenAI's model family with GPT-5.3-Codex.
Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.
No Continuum hosted id is verified for this model, so this page does not invent a curl target. Use the first-party API or the OpenRouter slug openai/gpt-5.4.