Models/DeepSeek/DeepSeek V4 Flash

DeepSeek V4 Flash

0731 is current Flash. It led OpenRouter This Week at 11.3T tokens. Continuum hosts deepseek-v4-flash. 0423 is lineage. Official DeepSeek USD for Flash is not the Pro table.

01

Identity

1,310,720Context in
393,216Context out
2026-07-31 (OpenRouter created)Released
OpenWeights
Canonical name
DeepSeek V4 Flash
Aliases
deepseek-v4-flash, deepseek-v4-flash-0731
Continuum hosted id
deepseek-v4-flash
OpenRouter slug
deepseek/deepseek-v4-flash-0731
Hugging Face
deepseek-ai/DeepSeek-V4-Flash-0731
Modalities
text
Open weights Hosted on Continuum Vendor lab

Current weights 0731. Unversioned deepseek/deepseek-v4-flash is 0423 lineage. Continuum hosted id is the bare deepseek-v4-flash.

HF config.json (fetched 19 Aug 2026): DeepseekV4ForCausalLM, 43 layers, hidden 4096, 256 routed experts, 6 experts/token, 1 shared expert, YaRN to 1,048,576, DSpark attached (dspark_block_size 5). Official DeepSeek docs list context 1M and max out 384K. OpenRouter’s 0731 row is 1,310,720 / 393,216. Those are two published windows; do not average them.

No Jinja chat template: the repo ships an encoding/ folder (same pattern as Pro 0813).

02

Should I use this for coding agents

Use Flash for volume, for hosted DeepSeek on Continuum, and when $0.10 DeepSWE tasks matter more than the top pass rate. Official board: 53% Pass@1 ±4% at max, $0.10 per task, 108k output, 153 steps.

DeepSWEmini-swe-agent
53%±4%
bench
DeepSWE
version
v1.1
split
public 113 tasks
harness
mini-swe-agent
effort
max
n
113
metric
Pass@1
value
53%
ci
±4%
$/task
$0.10
tokens
108k out
steps
153
independent as of 2026-08-13 source
Artificial Analysispublished integer
52
bench
Artificial Analysis
version
Intelligence Index
harness
published integer
metric
Intelligence Index
value
52
independent as of 2026-08-19 source
Vals Indexvals.ai card
53.57%
bench
Vals Index
version
hero index
harness
vals.ai card
metric
Vals Index
value
53.57%
independent as of 2026-08-19 source

Ranked view: Best coding models: the independent leaderboard puts this row and every other card that carries an independent coding score on one board, with the confidence intervals left visible.

03

When not to use it

Trust this before you buy

  • If you need the Pro checkpoint or a higher DeepSWE band, use V4 Pro 0813 (not hosted). Pro’s official USD table is a different column on the same DeepSeek page; do not paste those cents here.
  • If you need 70%+ on this board, use Opus, Sol, or Fable and pay closed-weight rates.
  • 153 steps on DeepSWE is a long agent. Cheap tokens still mean a long wall clock. Peak vs off-peak is a real official split. Do not paste Pro cents onto Flash. No official GGUF on this repo.

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.

04

Artifact and Hugging Face downloads

Artifact

SPDX / license
MIT
Total params
HF safetensors.total: 304,180,418,494 stored tensors. Do not invent an AA 1600B figure for Flash; that number lives on the Pro card.
Architecture
DeepseekV4ForCausalLM · 43 layers · 256 routed experts · 6 experts/tok · 1 shared · hidden 4096 · YaRN 1M · DSpark attached
Native precision
HF tensors I8 + F8_E4M3 + BF16 + F32 + I64. config expert_dtype fp4, quant_method fp8 e4m3 (block 128×128)
Files
48 safetensor shards (model-00001-of-00048 …) plus encoding/ and inference/
Repo size
HF API usedStorage 166,888,735,421 bytes (155 GiB)
HF created
2026-07-31T07:30:24Z
Sampling
Same family as Pro on the official docs: thinking (default) and non-thinking. Do not copy Pro’s vendor README sampling block onto Flash without a Flash README citation.
Chat template
No Jinja. Official encoding/ Python scripts + fixtures (OpenAI-compatible messages → string).
Paper
arXiv:2606.19348: DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence (HF tag on this repo)
Official HF repo
deepseek-ai/DeepSeek-V4-Flash-0731
HF downloads
2,330,940
HF likes
3,545

Preview / prior weights: Prior weights: DeepSeek-V4-Flash (0423 lineage).

Community quants: Any GGUF or AWQ you find under other orgs is community, not official.

DSpark speculative decoding ships in this checkpoint too (dspark_block_size 5). That is a config field, not a second repo.

Official DeepSeek API docs (fetched 19 Aug 2026) put Flash and Pro on one page with two price columns. Flash is the left column. Copying Pro cents onto Flash is a fabricated schedule.

Official weights

This is the model repo, not Download for Mac. Downloads and likes from the Hugging Face API on 2026-08-19.

Hub file tree

Safetensor shards
48
Shard bytes
166,886,535,336 bytes
Other file bytes
12,125,738 bytes
Tree file count
74
Hub usedStorage
166,888,735,421 bytes

usedStorage and the sum of .safetensors sizes can differ (LFS pointers, non-shard files, duplicate copies). Cited separately. Source: recursive tree API, 2026-08-19.

PathBytes
model-00001-of-00048.safetensors1,059,061,856
model-00002-of-00048.safetensors3,566,321,192
model-00003-of-00048.safetensors3,566,321,192
model-00004-of-00048.safetensors3,596,229,272
model-00005-of-00048.safetensors3,568,768,976
model-00006-of-00048.safetensors3,590,024,776
model-00007-of-00048.safetensors3,568,768,976
model-00008-of-00048.safetensors3,590,024,776
model-00009-of-00048.safetensors3,568,768,976
model-00010-of-00048.safetensors3,590,024,776
model-00011-of-00048.safetensors3,568,768,976
model-00012-of-00048.safetensors3,590,026,352

Showing the first 12 of 48 safetensor shards. The tree also holds 26 non-shard files, none of them listed here. The snapshot was capped at 12 shards when the tree API was read on 2026-08-19, so the remaining 36 are on the Files tab, not on this page. Totals above are the full tree API sum, not this slice. Byte sizes are Hub tree API size fields, never invented.

05

Continuum serving

Continuum hosts deepseek-v4-flash (19 Aug 2026 public allowlist).

First-party DeepSeek API model id on the official pricing page is deepseek-v4-flash, version DeepSeek-V4-Flash-0731. OpenRouter slug is deepseek/deepseek-v4-flash-0731.

Continuum hosted id deepseek-v4-flash

click to select

We list first-party and OpenRouter, plus Continuum only when the live public allowlist named the id. This is not a 15-host routing table.

06

Economics

DeepSWE official: $0.10 per task at max (108k out, 153 steps). That is the hero economics number.

First-party DeepSeek API docs, fetched 19 Aug 2026: one USD schedule with peak and off-peak. Peak hours 01:00–04:00 and 06:00–10:00 UTC; all other hours off-peak at half. Flash cache-hit / cache-miss / output: peak $0.014 / $0.44 / $1.32 per million; off-peak $0.007 / $0.22 / $0.66. Concurrency cap 2500 (Pro is 500). These are not the Pro column ($0.044 / $1.32 / $3.96 peak).

OpenRouter catalog 19 Aug 2026 for deepseek-v4-flash-0731 lists $0.14 / $0.28. Continuum’s hosted rate card matches those cents. That catalog pair is not the official DeepSeek Flash peak/off-peak table. Artificial Analysis printed the official Flash peak miss/out pair ($0.44 / $1.32). Cite all three; do not collapse them into one mystery number.

Price this model Opens the pricing calculator preloaded with DeepSeek V4 Flash.

Artificial Analysis Intelligence Index 52, fetched 19 Aug 2026. AA printed $0.44 / $1.32 for Flash: that is the official DeepSeek peak cache-miss / output pair, not a fourth table. AA did not publish a DeepSWE-style CI, step count, or $/task on the chip we will copy; those fields stay omitted on the AA tile. Do not average AA cost with DeepSWE $0.10/task.

07

Provenance

ClaimSourceAs of
DeepSWE 53% ±4% at maxDeepSWE official board2026-08-13
Official Flash peak $0.014 / $0.44 / $1.32 and off-peak $0.007 / $0.22 / $0.66; peak 01:00–04:00 and 06:00–10:00 UTC; Flash concurrency 2500DeepSeek API Models & Pricing2026-08-19
HF safetensors.total 304,180,418,494; usedStorage 166,888,735,421 bytes; 48 shards; likes 3,545; created 2026-07-31T07:30:24Z; arXiv:2606.19348Hugging Face API DeepSeek-V4-Flash-07312026-08-19
config DeepseekV4ForCausalLM, 43 layers, 256 routed experts, 6/tok, DSpark 5HF config.json DeepSeek-V4-Flash-07312026-08-19
OR catalog $0.14 / $0.28 for flash-0731; AA printed $0.44 / $1.32OpenRouter catalog + Artificial Analysis DeepSeek V4 Flash2026-08-19
Hosted id deepseek-v4-flashGET /v1/chat/hosted/models/public2026-08-19
11.3T week tokens on flash-0731OpenRouter rankings This Week2026-08-19
OpenRouter id deepseek/deepseek-v4-flash-0731, context 1,310,720OpenRouter /api/v1/models2026-08-19
HF downloads 2,330,940Hugging Face API deepseek-ai/DeepSeek-V4-Flash-07312026-08-19
48 safetensor shards; shard bytes 166,886,535,336; tree files 74Hugging Face tree API deepseek-ai/DeepSeek-V4-Flash-07312026-08-19
Spaces API returned 22; official collection DeepSeek-V4 (4 items, 833 upvotes)Hugging Face Spaces / collections API2026-08-19
08

Compare, FAQ, and Get Plus

Same official DeepSWE harness (mini-swe-agent, public 113 tasks), fetched 2026-08-19. Different effort labels are the lab's own setting on that board, shown here rather than normalized.

Continue through DeepSeek's model family with R1.

FAQ

Is the hosted id the 0731 build?

Continuum lists deepseek-v4-flash without a date suffix. Official DeepSeek docs name the current version DeepSeek-V4-Flash-0731. We do not invent a silent pin beyond that.

Are Flash and Pro the same price off-peak?

No. Official docs: Flash off-peak $0.007 / $0.22 / $0.66; Pro off-peak $0.022 / $0.66 / $1.98. Same page, two columns.

Why does OpenRouter say $0.14 / $0.28 if official peak is $0.44 / $1.32?

Those are different published tables. Official DeepSeek is peak/off-peak USD. OpenRouter’s 0731 catalog row is $0.14 / $0.28. AA printed official peak miss/out. We cite each source; we do not pick a winner.

Why does the lab blog disagree with DeepSWE or Scale?

Lab posts pick a harness, an effort, a split, and sometimes a private eval set. DeepSWE publishes the official mini-swe-agent row with cost, tokens, and steps. Scale SWE-bench Pro only counts when the public shared-harness board has a row we can fetch. A higher lab number is usually a different test, not a better one.

OpenAI-compatible call

Verified against the live public allowlist on 2026-08-19. Base URL is https://continuumcode.ai/v1. Keys are cont_sk_ from Settings, Account, Inference API. Personal keys need Plus or above.

curl https://continuumcode.ai/v1/chat/completions \
  -H "Authorization: Bearer $CONTINUUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek-v4-flash","messages":[{"role":"user","content":"Review this diff."}]}'

Plus is $25/mo with $25 weekly hosted usage. The Mac app stays free with your own keys. Get Plus is not Download for Mac.