DeepSeek V4 Flash Vision Pricing: $0.14/$0.28 per 1M
Daniel Okonkwo
Senior ML Engineer

TLDRDeepSeek V4 Flash Vision costs $0.14 input and $0.28 output per 1M tokens, images at up to 384 tokens each. Full cost breakdown, as of Aug 22, 2026.
How Much Does DeepSeek V4 Flash Vision Cost? $0.14/$0.28 per 1M Tokens With Images at 384 Tokens Each
DeepSeek V4 Flash Vision (deepseek-v4-flash-vision-exp) costs $0.14 per 1M input tokens and $0.28 per 1M output tokens, the same rate as text-only V4-Flash, with images billed as tokens at up to 384 tokens per image. It went live on the DeepSeek API Platform on August 21, 2026 as an experimental multimodal model. Cached input tokens drop to $0.0028 per 1M, and the Files API for uploading images is free to use. All figures below are as of August 22, 2026.
Key Takeaways
- Text pricing: $0.14 per 1M input tokens (cache miss), $0.28 per 1M output tokens — identical to DeepSeek V4-Flash.
- Cache hit input: $0.0028 per 1M tokens, a 50x discount versus cache-miss input.
- Image pricing: images are tokenized at up to 384 tokens each and billed at the text rate; no separate per-image fee reported.
- Files API: free to upload and reference images by
file_idacross requests. - Status: experimental (
-exp), API-only, no open weights yet. - Positioning: roughly one-third the input cost of DeepSeek V4-Pro, per community reports.
DeepSeek V4 Flash Vision Pricing Breakdown
DeepSeek V4 Flash Vision inherits V4-Flash pricing directly. The official release note states images are "tokenized for billing: up to 384 tokens each, at V4-Flash pricing," and DeepSeek's release announcement confirms the model matches V4-Flash on text capabilities. The per-token numbers come from DeepSeek's published Models & Pricing page for deepseek-v4-flash.
| Component | Price (per 1M tokens) | Source |
|---|---|---|
| Input tokens (cache miss) | $0.14 | DeepSeek Models & Pricing |
| Input tokens (cache hit) | $0.0028 | DeepSeek Models & Pricing |
| Output tokens | $0.28 | DeepSeek Models & Pricing |
| Image input | Tokenized at up to 384 tokens each, then text rate | DeepSeek release note |
| Files API (image upload) | Free | DeepSeek release note |
| Volume-based discount tiers | Not yet announced | — |
| Dedicated subscription bundle | Not yet announced | — |
The Reddit thread mirroring the announcement summarizes it plainly: "Same pricing with flash too, with max image token count of 384." That community read matches the official documentation.

Source: @OpenRouter
Note the model carries the -exp (experimental) suffix. DeepSeek states it "reserves the right to adjust" prices, so these figures are a snapshot, not a locked contract.
What That Costs in Practice
All estimates below use DeepSeek's published rates: $0.14 per 1M input, $0.28 per 1M output, and images at the 384-token maximum. Cache hits are not assumed.
| Scenario | Tokens | Math | Cost |
|---|---|---|---|
| One image + short question (384 img + 100 text in, 300 out) | 484 in / 300 out | (484 × $0.14 + 300 × $0.28) / 1M | ~$0.00016 |
| 1,000 images described (384 img + 50 text in, 200 out each) | 434K in / 200K out | (434K × $0.14 + 200K × $0.28) / 1M | ~$0.117 |
| Screenshot agent loop (5 images + 4K text in, 1K out) | 5,920 in / 1,000 out | (5,920 × $0.14 + 1,000 × $0.28) / 1M | ~$0.0011 |
| A single 384-token image, input only | 384 in | 384 × $0.14 / 1M | ~$0.0000538 |
A single image billed at its 384-token cap costs about $0.0000538 — roughly 5 cents per 1,000 images before any accompanying text. At the $0.28 per 1M output rate, a 1,000-token model response costs $0.00028. These are among the lowest published multimodal rates for a model of this class.
How DeepSeek V4 Flash Vision Pricing Compares
DeepSeek V4 Flash Vision undercuts its larger sibling on every axis. The comparison below uses only figures published in DeepSeek's own pricing docs.
| Model | Input (cache miss) | Output | Notes |
|---|---|---|---|
| DeepSeek V4 Flash Vision | $0.14 / 1M | $0.28 / 1M | Vision + text, images ≤384 tokens |
| DeepSeek V4 Flash | $0.14 / 1M | $0.28 / 1M | Text only |
| DeepSeek V4 Pro | $0.435 / 1M | $0.87 / 1M | Text only |
Input on V4 Flash Vision is roughly one-third of V4-Pro's $0.435 per 1M, and output is about a third of V4-Pro's $0.87 per 1M. According to one community report, V4-Flash-Vision-Exp costs "one-third of V4-Pro" — consistent with the published numbers. For deeper context on how the underlying Flash model reached this price point, see our earlier DeepSeek V4 Flash 0731 release analysis.
Free Tier, Limits, and Access
DeepSeek V4 Flash Vision is a paid API model with no announced free credit allowance. Billing works on DeepSeek's standard deduction rule: expense = tokens × price, drawn from your topped-up or granted balance.
What is free: the Files API. DeepSeek's release note confirms it is "free to use" — upload an image once, then reference it by file_id across requests to save bandwidth. Images can also be sent via base64 or external URL.
On limits, DeepSeek's pricing table lists a concurrency limit of 2,500 for V4-Flash. Community reports (single-source, unconfirmed) say V4-Flash-Vision-Exp inherits the same 2,500 concurrency ceiling, a 1M-token context window, and 384K maximum output. Treat those as pre-release community figures until DeepSeek publishes a dedicated spec table.
Access is currently through the official DeepSeek API Platform with model='deepseek-v4-flash-vision-exp', supporting Chat Completions, Messages, and Responses formats. DeepSeek Harness 0.1.1 ships with out-of-the-box support; see our DeepSeek Harness explainer for how that agent runner works. If you need a hosted multimodal chat endpoint today while DeepSeek's vision model remains experimental, comparable models such as Claude Opus 4.8 are available through the kie.ai API.
What We Don't Know Yet
DeepSeek has published the core token rates but left several pricing-adjacent details open:
- Peak vs. off-peak split: community reports describe peak hours (Beijing time 09:00–12:00 and 14:00–18:00) with higher CNY-denominated rates, but DeepSeek's English pricing page shows only the single USD figures above. Time-of-day pricing for the vision model is not officially confirmed.
- Volume discounts or committed-use tiers: not yet announced.
- A production (non-
exp) vision model and its price: not yet announced. The current model is explicitly experimental. - Open weights: none released; an informed community prediction suggests weights for a future final model may open, but that is unverified.
Frequently Asked Questions
How much does the DeepSeek V4 Flash Vision API cost?
DeepSeek V4 Flash Vision (deepseek-v4-flash-vision-exp) is billed at V4-Flash pricing: $0.14 per 1M input tokens on a cache miss, $0.0028 per 1M input tokens on a cache hit, and $0.28 per 1M output tokens, as of August 22, 2026.
How are images priced in DeepSeek V4 Flash Vision?
Images in DeepSeek V4 Flash Vision are tokenized for billing at up to 384 tokens each, then charged at the standard V4-Flash text rate. There is no separate per-image fee reported.
Is DeepSeek V4 Flash Vision free?
DeepSeek V4 Flash Vision is not free; it is a paid API model. However, the associated Files API used to upload and reference images is free to use, per DeepSeek's documentation.
Is DeepSeek V4 Flash Vision cheaper than DeepSeek V4 Pro?
Yes. DeepSeek V4 Flash Vision uses V4-Flash pricing at $0.14 input and $0.28 output per 1M tokens, roughly one-third the input cost of V4-Pro's $0.435 per 1M input and well below its $0.87 per 1M output.
How much does a single image cost with DeepSeek V4 Flash Vision?
A single image billed at the 384-token maximum costs about $0.0000538 at DeepSeek V4 Flash Vision's $0.14 per 1M input rate, or roughly 5 cents per 1,000 images before any accompanying text tokens.
Does DeepSeek V4 Flash Vision have a free tier or rate limits?
DeepSeek V4 Flash Vision has no announced free credit tier. It inherits V4-Flash's 2,500 concurrency limit according to community reports, and the Files API for image reuse is free.
Are DeepSeek V4 Flash Vision prices final?
DeepSeek V4 Flash Vision is labeled experimental (the -exp suffix), and DeepSeek states it reserves the right to adjust prices. The figures here are current as of August 22, 2026.
What to Watch Next
Three signals will change this page. First, whether DeepSeek graduates the model out of -exp and publishes a dedicated pricing row separate from V4-Flash. Second, whether official English docs confirm any peak/off-peak time-of-day pricing that community posts describe in CNY. Third, whether open weights for a final V4 vision model appear, which would add a self-hosting cost path alongside the API.
Building similar multimodal chat features? On kie.ai you can try Claude Opus 4.8, Gemini 3.8 Flash, and GPT-5.6.
About Daniel Okonkwo
Daniel writes about inference systems, model architecture, and what new releases actually change for builders.
View all posts by Daniel Okonkwo