- bench
- DeepSWE
- version
- v1.1
- split
- public 113 tasks
- harness
- mini-swe-agent
- effort
- xhigh
- n
- 113
- metric
- Pass@1
- value
- 57%
- ci
- ±3%
- $/task
- $3.73
- tokens
- 95k out
- steps
- 111
API Max SKU. 297B and #39 on /models Top Weekly. Qwen is 5.7% of text requests. DeepSWE lists Max at 57% ±3% / $3.73 per task at xhigh. AA Intelligence Index 58. vals hero 51.84% ±1.29. Official Alibaba / qwen.ai pages we opened were JS shells with no extractable dollar table. Continuum does not host it. Open 2.4T weights are a different identity.
Closed API. OpenRouter catalog created unix 1785731612 (3 Aug 2026). 1,000,000 in / 131,072 out; text + image + video in, text out. hugging_face_id empty. Open sibling Qwen/Qwen3.8-2.4T-A95B is lineage on the hub, not this card’s download.
Use Max when you want Qwen’s current API flagship. DeepSWE official: 57% ±3% at xhigh, $3.73 per task, 95k out, 111 steps.
vals.ai Qwen3.8 Max card, fetched 19 Aug 2026: hero Index 51.84% ±1.29, $3.997/test, 90 min 46 s latency. Accuracy Rankings bars printed 0.0% placeholders: omitted. No Updates prose with named non-zero benches was present on that HTML.
Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.
Copied from AA or vals HTML we opened. 0.0% placeholder bars are omitted. Do not average these into the hero tiles, and do not invent CI, $/task, or steps for Artificial Analysis.
| Bench | Printed | Note | Source | URL | As of |
|---|---|---|---|---|---|
| AA list pair | $2.00 / $6.00 per 1M | Printed AA pricing sentence. Matches OpenRouter. | Artificial Analysis Qwen3.8 Max | source | 2026-08-19 |
| AA throughput | 44.4 tok/s | Printed Speed row. Not a coding score. | Artificial Analysis Qwen3.8 Max | source | 2026-08-19 |
| AA Index eval cost | $1741.41 | Printed “it cost $1741.41 to evaluate Qwen3.8 Max on the Intelligence Index.” Not DeepSWE $/task. | Artificial Analysis Qwen3.8 Max | source | 2026-08-19 |
| vals Index extras | $3.997/test · 90 min 46 s | Printed Cost / Test and Latency on the vals hero. Hero Accuracy is 51.84% ±1.29. | vals.ai Qwen3.8 Max | source | 2026-08-19 |
Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.
Qwen3.8 Max has no HF id on OpenRouter. The open 2.4T repo is a sibling, labeled lineage, not a fake Max download.
Lab card and OpenRouter listing only. No Hugging Face Files button, because there is no official repo to point at.
Fetched OpenRouter catalog on 2026-08-19. hugging_face_id was empty.
Not listed as a Continuum hosted id. OpenRouter qwen/qwen3.8-max and first-party Qwen.
We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.
DeepSWE official: $3.73 per task at xhigh (95k out, 111 steps). That is the hero economics number.
OpenRouter catalog 19 Aug 2026: $2 / $6 per million, cache read $0.25 /M, cache write $2.50 /M, created 1785731612. No official Alibaba per-M table extracted from the JS shell we fetched.
Artificial Analysis Qwen3.8 Max page (19 Aug 2026) prints the same $2.00 / $6.00 pair, 44.4 tok/s, and “it cost $1741.41 to evaluate Qwen3.8 Max on the Intelligence Index.” vals prints $3.997 per Index test and 90 min 46 s.
Artificial Analysis Intelligence Index 58, fetched 19 Aug 2026. Same page prints $2.00 / $6.00, 44.4 tok/s, and $1741.41 to run the Index. The AA tile is the published integer 58. AA did not print a DeepSWE-style CI or step count we will copy onto the chip. Do not average AA cost with DeepSWE $3.73/task.
| Claim | Source | As of |
|---|---|---|
| DeepSWE 57% ±3% at xhigh | DeepSWE official board | 2026-08-13 |
| OR catalog $2 / $6, cache read $0.25/M, cache write $2.50/M, 1,000,000 / 131,072, created 1785731612, hugging_face_id empty | OpenRouter catalog | 2026-08-19 |
| Official Alibaba/Qwen pages opened; JS shells; no extractable $ table in saved HTML | qwen.ai | 2026-08-19 |
| AA Index 58; $2.00 / $6.00; 44.4 tok/s; $1741.41 to evaluate Index | Artificial Analysis Qwen3.8 Max | 2026-08-19 |
| vals hero 51.84% ±1.29; $3.997/test; 90 min 46 s; 0.0% Accuracy Rankings bars omitted | vals.ai Qwen3.8 Max | 2026-08-19 |
| OpenRouter id qwen/qwen3.8-max, context 1,000,000 | OpenRouter /api/v1/models | 2026-08-19 |
Same official DeepSWE harness (mini-swe-agent, public 113 tasks), fetched 2026-08-19. Different effort labels are the lab's own setting on that board, shown here rather than normalized.
Continue through Qwen's model family with Qwen3.7 Plus.
No. Qwen/Qwen3.8-2.4T-A95B is lineage on the hub. This card is the closed Max API. No fake Max download.
qwen.ai / Alibaba pages we opened were JS shells. No extractable per-M table. OpenRouter $2 / $6 plus cache $0.25 / $2.50 write is the sourced list.
Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.
No Continuum hosted id is verified for this model, so this page does not invent a curl target. Use the first-party API or the OpenRouter slug qwen/qwen3.8-max.