DeepSeek V4 Flash lost all 9 benchmarks to Opus 4.8
V4 Flash costs $0.14 in and $0.28 out per million tokens. I checked whether its 97 to 99 percent discount makes DeepSeek's benchmark losses worth it.
On this page
- How close is DeepSeek V4 Flash to Opus 4.8?
- What is DeepSeek V4 Flash 0731?
- Which Opus 4.8 score should you believe?
- What do the first independent numbers show?
- What does $0.28 per million tokens buy in practice?
- What do you give up for the discount?
- My verdict: test V4 Flash as a worker, not the reviewer
DeepSeek V4 Flash is not an Opus 4.8 replacement. DeepSeek’s own launch table puts Opus ahead on all nine benchmarks, although two gaps are small [2]. The reason to care is price: V4 Flash costs $0.14 per million input tokens and $0.28 per million output [3], against $5 and $25 for Opus 4.8 [4].
How close is DeepSeek V4 Flash to Opus 4.8?
DeepSeek V4 Flash is close to Opus 4.8 on two of DeepSeek’s nine tests, but it is well behind on the repository-heavy benchmarks. Opus leads by 0.5 points on Agents’ Last Exam and 2.3 on Terminal-Bench 2.1, then opens gaps of 15.5 points on NL2Repo and 12.1 on DSBench-Hard [2].
| Benchmark | V4 Flash 0731 | Opus 4.8 | GLM-5.2 |
|---|---|---|---|
| Terminal-Bench 2.1 | 82.7 | 85.0 (Best value in this row) | 81.0 |
| NL2Repo | 54.2 | 69.7 (Best value in this row) | 48.9 |
| Cybergym | 76.7 | 83.1 (Best value in this row) | - |
| DeepSWE | 54.4 | 58.0 (Best value in this row) | 46.2 |
| Toolathlon Verified | 70.3 | 76.2 (Best value in this row) | 59.9 |
| Agents' Last Exam | 25.2 | 25.7 (Best value in this row) | 23.8 |
| AutomationBench Public | 25.1 | 27.2 (Best value in this row) | 12.9 |
| DSBench FullStack | 68.7 | 71.6 (Best value in this row) | 61.8 |
| DSBench Hard | 59.6 | 71.7 (Best value in this row) | 54.5 |
The pattern matters more than any one row. V4 Flash stays near Opus on terminal-style tasks, but falls further behind when the job involves understanding or building a repository. It still leads GLM-5.2 on all eight rows the two models share.
DeepSeek makes a narrower claim than much of the coverage. The model card says the new build outperforms the larger V4 Pro preview while activating far fewer parameters [2]. Publishing a chart that shows every loss against Opus is also more useful than only publishing the rows a new model wins.
The comparison is already dated, however. Opus 4.8 was Anthropic’s strongest model when it launched on 28 May 2026 [7], but ten weeks later it appears in the legacy section of Anthropic’s model list [8]. Opus 5 arrived on 24 July at the same $5 input and $25 output rates, and Anthropic now points new work at it, with Fable 5 above both [9]. V4 Flash is competing with Anthropic’s May frontier, not its August one.
What is DeepSeek V4 Flash 0731?
DeepSeek-V4-Flash-0731 shipped on 31 July 2026 as the official, re-post-trained release of the V4 Flash preview [1]. The preview had been public since 24 April. DeepSeek kept the architecture and size, then redid only the final training stages. The result is a mixture-of-experts model with 284 billion total parameters, 13 billion active per token and a one-million-token context window [5].
The weights are on Hugging Face under an MIT license, so commercial use and self-hosting are unrestricted [2].
On the API side, DeepSeek swapped the build in place: the model string is still deepseek-v4-flash, and existing integrations moved to the new version without a config change [1]. That is convenient for users and awkward for anyone citing results, because “V4 Flash” now names two different builds depending on the date a number was measured. The model exposes a reasoning effort parameter with three levels, low, high and max, instead of shipping a separate reasoning variant [2].
The release attracted attention because of the price. Nikkei Asia described it as part of an escalating price war in AI models and reported that DeepSeek plans to introduce peak-hour pricing later [6]. DeepSeek’s pricing page confirms that prices will double during peak hours once the policy takes effect, which had not happened by 3 August 2026 [3]. The larger V4 Pro remains in preview at $0.435 in and $0.87 out, with its official release announced as coming soon [1] [3].
- input price, cache miss
- 0.14 $/MTok
- Opus 4.8: $5
- output price
- 0.28 $/MTok
- Opus 4.8: $25
- context window
- 1M
- max output 384K tokens
- active parameters
- 13B
- of 284B total, MIT license
Per million output tokens, Opus 4.8 costs about 89 times as much as V4 Flash. That difference is large enough to matter, but a price comparison still needs a credible capability baseline. The Opus benchmark scores are where that becomes difficult.
Which Opus 4.8 score should you believe?
No single Opus 4.8 score is stable enough to anchor this comparison. DeepSeek’s harness reports 85.0 on Terminal-Bench 2.1 [2], the public Terminal-Bench leaderboard lists 78.9 percent with Claude Code [10], and Anthropic’s system card reports 74.6 percent at high effort on the Terminus-2 harness [7].
Show the data as a table
| Source | Value |
|---|---|
| DeepSeek Harness, max effort | 85.0% |
| Terminal-Bench leaderboard, Claude Code | 78.9% |
| Anthropic system card, Terminus-2, high effort | 74.6% |
V4 Flash’s own 82.7 lands inside that 10.4-point spread. Depending on which Opus result you use, DeepSeek’s model is 2.3 points behind, 3.8 ahead or 8.1 ahead. The model scores alone cannot tell us which one is better at terminal tasks.
DeepSeek reported the highest Opus score of the three, so its result does not lowball the competitor. The 10.4-point range comes from different agent harnesses, effort settings and timeouts. Anthropic’s system card also notes that Terminal-Bench penalizes slow endpoints because tasks run against fixed wall-clock limits [7]. These setups cannot be compared directly.
A third party running both models through one harness could settle the comparison. I could not find a public release of the DeepSeek Harness or an independent reproduction of any of the nine rows as of 3 August 2026. Until one exists, this remains a vendor result rather than an independently reproduced comparison.
What do the first independent numbers show?
Artificial Analysis places V4 Flash at the top of the budget class rather than the frontier. Its Intelligence Index score of 50 is ten points above the April preview and one point behind GPT-5.6 Luna at max effort [11]. Opus 4.8 is not in that index, so this does not reproduce DeepSeek’s comparison.
On the same leaderboard as of 3 August 2026, Gemini 3.6 Flash also scored 50, GLM-5.2 reached 51, GPT-5.6 Terra reached 55 and Kimi K3 reached 57 [12]. That is the only standardized independent placement I found for the 0731 build.
Artificial Analysis also confirmed DeepSeek’s central claim that V4 Flash beats the larger V4 Pro at a lower cost. Running the full index cost $72.02 with V4 Flash [13], compared with $176.34 for V4 Pro, which scored six points lower at 44 [14]. That is a better result at about 40 percent of the cost.
Token use still matters to the total bill. V4 Flash needed roughly 206 million output tokens to complete the index, 12 percent fewer than the April build [11], but more than the 180 million V4 Pro used [14]. The model is cheap per token rather than economical with tokens, so estimates based on another model’s token counts will run low.
Two other Artificial Analysis tests show a large improvement over the April preview. On GDPval-AA v2, a work-task comparison scored by preference votes, the new build reached an Elo of 1559, up from 1189 [11]. Its hallucination rate on AA-Omniscience improved by 12 points to 84 percent, which is progress from a bad starting point rather than a good result [11].
The WebDev leaderboard also places V4 Flash below the frontier models. On Arena, where people vote on generated web apps, V4 Flash at high effort held seventh place with an Elo of 1577 on 3 August 2026, compared with 1705 for Opus 5 Max at the top [15].
Speed remains unknown for this build. I found no independent tokens-per-second measurement for 0731, so any speed figure in circulation on 3 August 2026 belonged to an older V4 Flash build [16].
What does $0.28 per million tokens buy in practice?
The early reports describe a cheap coding worker that performs best with close supervision, not a general replacement for a stronger model. The actual bill can be lower than the headline price when agent workflows reuse cached context, while future peak pricing and reseller discounts make $0.28 only the starting point.
Cached input costs $0.0028 per million tokens, one fiftieth of the cache-miss rate [3], and agent workflows resend mostly unchanged context on every turn. In the other direction, DeepSeek says prices will double during Beijing peak hours once that policy activates [3]. At launch, OpenRouter added a 38 percent discount, roughly $0.09 in and $0.17 out [17]. The $0.17 or $0.18 output prices quoted elsewhere therefore refer to the reseller promotion, not DeepSeek’s list.
One caveat applies to every cross-vendor token price in this article. Anthropic’s pricing documentation states that Claude models from 4.7 onward use a tokenizer that produces roughly 30 percent more tokens for the same text [4], so per-token ratios between vendors are approximations. A factor of 89 on output survives a 30 percent correction with room to spare.
One Hacker News engineer posted a 30-day dashboard showing $4.55 for 3,467 API requests and 323 million tokens, with the verdict that the model “has not disappointed at all for coding or review tasks. For everything else, use another model” [18]. That averages about $0.014 per million tokens, below every listed rate except the cache hit, so most of the traffic must have been cached input. The same user says the model is weak at prose, improves with precise instructions and self-review tooling, and “you must still review 100% of the code as if it’s trying to sell you insurance” [19].
Another commenter hit a tight output-token limit through OpenRouter after the model spent too long reasoning, but still concluded that it “replaced the Kimi models for me” [20]. Simon Willison’s same-day test points to one useful setting: his default-effort result was clearly worse than the result at high [21]. Taken together, the reports support high effort, well-scoped tasks and a stronger model or human in the review seat.
What do you give up for the discount?
V4 Flash has no image input, and one controlled study found much stronger censorship on China-sensitive questions than on matched control questions [11] [22]. Those limits matter for multimodal coding workflows and any product that answers questions about politically sensitive subjects.
CTGT, a research firm that tests model behavior, measured a 45.45-point censorship gap between China-sensitive questions and structurally identical control questions. The researchers tested self-hosted weights to rule out filtering by the API provider, while a model distilled from V4 Flash outputs showed no significant trace of the same behavior [22]. The result therefore applies to the weights, not only to DeepSeek’s hosted endpoint.
Without image input, the model cannot run the screenshot-driven frontend loops used by multimodal agents [11]. The official API is optional because the MIT weights [2] are already available through OpenRouter and other Western providers [17]. Choosing another provider can change the endpoint, data jurisdiction and pricing, but it cannot remove behavior carried in the weights.
My verdict: test V4 Flash as a worker, not the reviewer
V4 Flash is worth testing because it delivers a model in this class at one to three percent of Opus rates, not because it beats Claude. DeepSeek’s own table says it loses, and the independent data places it with strong budget models rather than at the frontier.
My next step is to point V4 Flash at my usual test repository with reasoning effort at high and a stronger model in the review seat. If an independent harness eventually runs it beside Opus under the same conditions, I will revisit the comparison. Until then, the discount expands the tasks I can afford to automate, but it does not remove the need for review.
Sources
- Change log
- DeepSeek-V4-Flash-0731 model card
- Models & Pricing
- Pricing
- DeepSeek-V4 model card
- DeepSeek releases beta version of V4 models as AI price war heats up
- System Card: Claude Opus 4.8
- Models overview
- Introducing Claude Opus 5
- Terminal-Bench 2.1 leaderboard
- DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index
- Models leaderboard
- DeepSeek V4 Flash 0731: intelligence, performance and price
- DeepSeek V4 Pro: intelligence, performance and price
- Arena leaderboard
- DeepSeek V4 Flash providers
- DeepSeek V4 Flash
- Comment with 30-day usage figures on the V4 Flash release thread
- Comment on working practices with V4 Flash
- Comment on output limits
- deepseek-ai/DeepSeek-V4-Flash-0731
- Distillation censorship transfer