What Is DeepSeek V4 Flash? Specs, Price, Access
Marcus Bell
Frontier Models Correspondent

TLDRDeepSeek V4 Flash is a 284B MoE model with a 1M context and $0.28/M output pricing — beating V4-Pro Preview on 9 agent evals.
Inside DeepSeek V4 Flash: The $0.28 Output Price That Undercuts Opus 4.8 by 89x
DeepSeek V4 Flash is an open-weights Mixture-of-Experts language model developed by DeepSeek AI, released as the deepseek-v4-flash-0731 checkpoint on July 31, 2026 and now in public beta on the official DeepSeek API. It has 284 billion total parameters with 13 billion active per token, a 1 million token context window, and lists at $0.14 per 1M input tokens and $0.28 per 1M output tokens — pricing that Artificial Analysis measured as roughly 105x cheaper per completed benchmark task than Claude Fable 5. Same model architecture as the April 2026 V4 Flash Preview; the reported gains come entirely from a fresh post-training pass.
Key Takeaways
- DeepSeek V4 Flash 0731 launched July 31, 2026 as an official public beta, replacing the earlier preview checkpoint at the same model string.
- It is a sparse Mixture-of-Experts model: 284B total parameters, 13B active per token, 256 routed experts, 6 active per token.
- Pricing is $0.14/M input, $0.0028/M cached input, $0.28/M output — a 98% cache discount versus the ~90% industry norm.
- DeepSeek reports V4 Flash beating its own larger V4-Pro-Preview across nine agent evaluations, including Terminal Bench 2.1 at 82.7.
- Weights are MIT-licensed and published on Hugging Face; the model has been run locally on Apple M3 Ultra, dual DGX Spark, and 2x H200 rigs within days of release.
- Third-party Artificial Analysis Intelligence Index score is 50, one point behind GPT-5.6 Luna at 51.
What Is DeepSeek V4 Flash?
DeepSeek V4 Flash is the smaller, cheaper tier of DeepSeek's V4 model family, aimed at coding, tool use, and agent workflows where cost-per-task matters more than absolute peak intelligence. The deepseek-v4-flash-0731 release keeps the exact same architecture and parameter count as the earlier V4 Flash Preview from April 2026 — only the post-training pipeline was rerun. According to DeepSeek's own changelog, that post-training pass lifted agent capability past the larger V4-Pro-Preview across every benchmark they published.
The model is text-only. There is no vision, audio, or video input; that gap has been noted in community discussion as one of the clearest reasons a shop might still reach for GPT-5.6 Luna or Gemini instead. It supports a thinking mode with three reasoning-effort levels (low, high, max), tool calling, JSON output, and both OpenAI and Anthropic API formats, with native support for the Responses API. Our earlier reporting on the wider V4 launch context covers how Flash and Pro were staged.
The release status is unambiguous: V4 Flash 0731 is live in public beta on api.deepseek.com. V4 Pro, the larger sibling, remains on its preview endpoint with an official release described in DeepSeek's own materials as "coming as soon as possible."
DeepSeek V4 Flash at a Glance
| Attribute | Value |
|---|---|
| Developer | DeepSeek AI |
| Type | Sparse Mixture-of-Experts (MoE), reasoning-capable |
| Model ID | deepseek-v4-flash (points to DeepSeek-V4-Flash-0731) |
| Total parameters | 284B |
| Active parameters per token | 13B |
| Experts | 256 routed, 6 active per token |
| Modality | Text in, text out |
| Context window | 1,000,000 tokens |
| Max output | 384,000 tokens |
| Thinking mode | Supported; effort levels low / high / max |
| Input price (cache miss) | $0.14 per 1M tokens |
| Input price (cache hit) | $0.0028 per 1M tokens |
| Output price | $0.28 per 1M tokens |
| Concurrency limit | 2,500 |
| License | MIT |
| Weights | Published on Hugging Face |
| Release date | July 31, 2026 (public beta) |
| API formats | OpenAI, Anthropic, Responses API, Codex-adapted |
How DeepSeek V4 Flash Works
The design point is efficiency, not raw scale. V4 Flash activates only 13B of its 284B parameters for any given token, using a sparse MoE routing scheme with 256 experts and 6 chosen per token. That is why the model runs on consumer-grade multi-GPU rigs while producing scores comparable to models several times its size.
Two coined mechanics are worth naming, because both show up repeatedly in third-party writeups:
- DSpark speculative decoding. The Flash weights ship with an embedded DSpark draft module. Runtimes such as vLLM enable it with a single flag, and community benchmarks show meaningful throughput lifts — one MLX-based test reported decode speed rising from 25.5 to 47.4 tokens per second, an 86% gain, on 4K coding prompts.
- 98% cache discount. DeepSeek prices cached input at $0.0028 per million tokens, a 50x reduction from the cache-miss rate. Artificial Analysis called this out specifically: most vendors give roughly 90% off on cache hits, so the extra step from 90% to 98% is where much of Flash's per-task cost advantage comes from for repeated-context workloads such as coding agents.
The other efficiency lever is the KV cache itself. One community engineer measured V4 Flash needing 2.89 GB of KV cache memory for a 1M-token context, versus 29.2 GB for Kimi K3 at the same context length — roughly a 10x reduction, driven by architectural choices rather than by hardware supply.
Because architecture and parameter count are identical to the April preview, DeepSeek's own materials are explicit that the intelligence gains came from post-training. That has one direct engineering consequence: any prompt-brittle agent scaffolding built against the preview should be re-evaluated, because tool-calling style, refusal behavior, and verbosity are exactly what post-training tends to shift.
What You Can Do With DeepSeek V4 Flash
The community deployment pattern that emerged within 48 hours of release is a good map of intended use:
- Coding agents at high volume. OpenCode Go and Cline both added V4 Flash as a supported backend on day one. Real-world usage reports on r/GithubCopilot converged around $0.30 to $1 per day of intensive coding, even with "high" reasoning effort selected.

Source: @xiaohu
-
Terminal and shell agents. Terminal Bench 2.1 at 82.7 is DeepSeek's flagship reported score, and it landed within 2.3 points of Claude Opus 4.8 on the same public benchmark.
-
Local inference on prosumer hardware. Ivan Fioravanti demonstrated V4 Flash 0731 running on an M3 Ultra at ~27 tokens/s with mxfp4 quantization, later reaching ~37 tokens/s with the DwarfStar runtime. Unsloth published GGUF quantizations shortly after weights dropped.
-
High-volume translation and batch text work. One user reported cutting a 20–60 minute translation job to 3 minutes at under $0.10 per article — the sort of workload that only becomes economical once per-token pricing drops an order of magnitude.
-
Frontend code generation. V4-Flash-High scored 1,586 points on Frontend Code Arena, placing seventh overall and third among open-weights models.
If your workflow is chat rather than agents, you can also compare against Anthropic's chat lineup on kie.ai via Claude Sonnet 5, which sits in a similar general-purpose slot but with multimodal input.
How DeepSeek V4 Flash Compares
Two comparisons matter most: against the frontier proprietary models it undercuts on price, and against its own preview.
| Model | Terminal Bench 2.1 | AA Intelligence Index | Input $/1M | Output $/1M |
|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | 82.7 (vendor) / 79 (AA) | 50 | $0.14 | $0.28 |
| DeepSeek V4 Flash Preview | 61.8 | 40 | $0.14 | $0.28 |
| DeepSeek V4 Pro Preview | 72.1 | 44 | $0.435 | $0.87 |
| Claude Opus 4.8 | 85.0 | 56 | $5.00 | $25.00 |
| GLM-5.2 | 81.0 | 51 | Not yet confirmed | Not yet confirmed |
Two honest notes on the table above. First, DeepSeek self-reported Terminal Bench at 82.7 while Artificial Analysis measured 79 using its own harness — a 3.7-point gap that the vendor's methodology footnote (a DeepSeek Harness "minimal mode" not yet publicly released) helps explain but not eliminate. Agent benchmarks are highly harness-sensitive; treat vendor numbers as an upper bound until independently reproduced. Second, the same third-party evaluator placed V4 Flash one point below GLM-5.2 on its aggregate Intelligence Index, whereas DeepSeek's own comparison table shows Flash ahead of GLM on eight of nine agent tasks. Both can be true — different benchmarks, different question mixes.
For a longer look at the open-weights competitive picture, see our earlier writeup on Kimi K3 vs Claude.
Availability: How to Access DeepSeek V4 Flash
Three official access paths are live today:
- Direct API.
https://api.deepseek.com(OpenAI format) orhttps://api.deepseek.com/anthropic(Anthropic format). Set model todeepseek-v4-flashand you receive the 0731 checkpoint automatically. Documented on api-docs.deepseek.com. - Codex-adapted. DeepSeek shipped an official Codex integration guide, including a preconfigured system prompt. Existing Codex configurations can point at DeepSeek without a translation shim.
- Self-hosted weights. MIT-licensed weights are on Hugging Face as
deepseek-ai/DeepSeek-V4-Flash-0731. Community-maintained GGUF quantizations from Unsloth landed within hours, and runtimes including vLLM, LM Studio, and MLX-based stacks added support inside the first week.
A peak-hour pricing policy is announced in the official docs — during peak hours (9:00–12:00 and 14:00–18:00 Beijing time), prices will be 2x the listed rates. The effective date has not been published.
What We Don't Know Yet
- V4 Pro release date and final pricing. DeepSeek promises "as soon as possible." Multiple community voices expect a step-up in price-performance, but no dated announcement exists.
- Full DeepSeek Harness release. The agent evaluation harness DeepSeek used for its self-reported benchmarks is announced as forthcoming; independent reproduction of the vendor scores is not yet possible.
- Peak-hour pricing effective date. Announced but not scheduled.
- Safety and jailbreak behavior. Two low-follower community accounts report the 0731 checkpoint being easy to jailbreak; there is no independent verified writeup and no official response.
- Multimodal roadmap. V4 Flash is text-only. Whether a multimodal variant is planned inside the V4 family has not been stated.
Frequently Asked Questions
What is DeepSeek V4 Flash?
DeepSeek V4 Flash is an open-weights Mixture-of-Experts language model developed by DeepSeek AI, released July 31, 2026 as the V4-Flash-0731 checkpoint. It has 284B total parameters with 13B active per token, a 1M-token context window, and is priced at $0.14 per million input tokens and $0.28 per million output tokens.
How much does DeepSeek V4 Flash cost?
DeepSeek V4 Flash costs $0.14 per 1M input tokens on cache miss, $0.0028 per 1M input tokens on cache hit, and $0.28 per 1M output tokens through the official DeepSeek API. That is roughly 1/89th the output price of Claude Opus 4.8 at $25 per million.
Is DeepSeek V4 Flash open source?
Yes. DeepSeek V4 Flash weights are published on Hugging Face under the MIT License, meaning the model can be downloaded, self-hosted, quantized, and used commercially. The API service and the training data are not open.
What is the context window of DeepSeek V4 Flash?
DeepSeek V4 Flash supports a 1 million token context window and up to 384K output tokens per response, per the official DeepSeek API documentation. It supports thinking mode with adjustable reasoning effort levels of low, high, and max.
How does DeepSeek V4 Flash compare to Claude Opus 4.8?
DeepSeek V4 Flash trails Opus 4.8 on most published agent benchmarks — for example 82.7 vs 85.0 on Terminal Bench 2.1 — while costing roughly 89x less per output token. Independent tester Artificial Analysis reports DeepSeek completing equivalent benchmark tasks at approximately 105x lower total cost.
Can I run DeepSeek V4 Flash locally?
Yes. Community testers have run DeepSeek V4 Flash 0731 on Apple M3 Ultra hardware at around 27–37 tokens per second, and Unsloth has published GGUF quantizations that fit in roughly 156–162 GB of memory. Because most weights ship in FP4 with the rest in FP8, further aggressive quantization yields limited additional savings.
What is the difference between DeepSeek V4 Flash and V4 Pro?
V4 Flash is the smaller, cheaper tier at 284B total parameters with 13B active; V4 Pro is the larger 1.6T-parameter model with 49B active per token. According to DeepSeek's own reported scores, V4-Flash-0731 exceeds the V4-Pro-Preview across nine agent benchmarks despite being smaller, though the official V4 Pro release is still pending.
What to Watch Next
Three concrete signals will determine whether V4 Flash 0731's launch positioning holds up. First, the official DeepSeek Harness release — until then, the gap between vendor-reported and third-party benchmark scores stays unresolved. Second, the V4 Pro official launch, which DeepSeek has telegraphed but not dated; if the same post-training gains carry over, the price-performance frontier shifts again. Third, DeepSeek's capacity story: multiple community reports already flag occasional slowdowns during Chinese business hours, and a peak-hour 2x pricing policy is on the books without an effective date.
Building similar high-context coding and agent workflows? On kie.ai you can try Claude Opus 5, GPT-5.6, and Gemini 3.8 Flash.
About Marcus Bell
Marcus reports on frontier model launches and leaks, weighing community testing against official specs.
View all posts by Marcus Bell