- bench
- DeepSWE
- version
- v1.1
- split
- public 113 tasks
- harness
- mini-swe-agent
- effort
- xhigh
- n
- 113
- metric
- Pass@1
- value
- 55%
- ci
- ±2%
- $/task
- $3.70
- tokens
- 99k out
- steps
- 101
Meta’s Spark SKU. DeepSWE lists it at 55% ±2% / $3.70 per task at xhigh. AA title is “Muse Spark 1.2 (xhigh)”; Intelligence Index 57. Official Meta pages we opened printed no extractable per-million table. The live Continuum public allowlist we fetched does not host it. vals 404.
OpenRouter hugging_face_id is empty. Catalog created unix 1785959287 (5 Aug 2026). 1,048,576 in; max out unpublished on that catalog row. The Mac bundled catalog names muse-spark-1.2; the live public allowlist on 19 Aug 2026 does not. We say not listed as a live Continuum hosted id.
Use Spark when you want Meta’s current API SKU and you accept first-party or OpenRouter billing. DeepSWE official: 55% ±2% at xhigh, $3.70 per task, 99k out, 101 steps.
We did not open a vals.ai card for this identity. Chip omitted.
vals 404 as of 19 Aug 2026. Chip omitted.
Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.
Copied from AA or vals HTML we opened. 0.0% placeholder bars are omitted. Do not average these into the hero tiles, and do not invent CI, $/task, or steps for Artificial Analysis.
| Bench | Printed | Note | Source | URL | As of |
|---|---|---|---|---|---|
| AA list pair (xhigh tile) | $1.25 / $4.25 per 1M | Printed AA pricing sentence. Matches OpenRouter. | Artificial Analysis Muse Spark 1.2 | source | 2026-08-19 |
| AA Index eval cost | $639.27 | Printed “it cost $639.27 to evaluate Muse Spark 1.2 (xhigh) on the Intelligence Index.” Not DeepSWE $/task. | Artificial Analysis Muse Spark 1.2 | source | 2026-08-19 |
Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.
No official Hugging Face id on OpenRouter. No fake weights button. Llama 4 repos are different identities.
Lab card and OpenRouter listing only. No Hugging Face Files button, because there is no official repo to point at.
Fetched OpenRouter catalog on 2026-08-19. hugging_face_id was empty.
Not listed as a live Continuum hosted id on 19 Aug 2026. The Mac bundled catalog still names it; we will not pretend the public gateway serves it.
We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.
DeepSWE official: $3.70 per task at xhigh (99k out, 101 steps). That is the hero economics number.
OpenRouter catalog 19 Aug 2026: $1.25 / $4.25 per million, cache read $0.15 /M, web search $0.0025, created 1785959287. Continuum’s catalog rate card matches $1.25 / $4.25. No official Meta list extracted.
Artificial Analysis title is “Muse Spark 1.2 (xhigh)”. That page prints the same $1.25 / $4.25 pair and “it cost $639.27 to evaluate Muse Spark 1.2 (xhigh) on the Intelligence Index.” No printed tok/s sentence on that HTML besides the generic meta description.
Artificial Analysis title is “Muse Spark 1.2 (xhigh)”. Intelligence Index 57, fetched 19 Aug 2026. Same page prints $1.25 / $4.25 and $639.27 to run the Index. The AA tile is the published integer 57. AA did not print a DeepSWE-style CI or step count we will copy onto the chip. Do not average AA cost with DeepSWE $3.70/task.
| Claim | Source | As of |
|---|---|---|
| DeepSWE 55% ±2% at xhigh | DeepSWE official board | 2026-08-13 |
| OR catalog $1.25 / $4.25, cache read $0.15/M, search $0.0025, 1,048,576 in, max out unpublished, created 1785959287, hugging_face_id empty | OpenRouter catalog | 2026-08-19 |
| Official Meta pages opened; no extractable per-M table; llama.com / llama.meta.com docs 404 | Meta AI | 2026-08-19 |
| Not on live public allowlist | GET /v1/chat/hosted/models/public | 2026-08-19 |
| AA title Muse Spark 1.2 (xhigh); Index 57; $1.25 / $4.25; $639.27 to evaluate Index | Artificial Analysis Muse Spark 1.2 | 2026-08-19 |
| vals 404 | vals.ai Muse Spark 1.2 | 2026-08-19 |
| OpenRouter id meta/muse-spark-1.2, context 1,048,576 | OpenRouter /api/v1/models | 2026-08-19 |
Same official DeepSWE harness (mini-swe-agent, public 113 tasks), fetched 2026-08-19. Different effort labels are the lab's own setting on that board, shown here rather than normalized.
Continue through Meta's model family with Llama 4 Maverick.
Bundled catalogs can lead the public gateway. This card follows the live allowlist so we never invent a curl target.
ai.meta.com printed no extractable per-M table. llama.com / llama.meta.com docs 404. OpenRouter $1.25 / $4.25 is the sourced list.
Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.
No Continuum hosted id is verified for this model, so this page does not invent a curl target. Use the first-party API or the OpenRouter slug meta/muse-spark-1.2.