Releases

DeepSeek compared V4 Flash to Opus 4.8 and lost every row

DeepSeek's launch table for V4 Flash shows Claude Opus 4.8 ahead on all nine benchmarks. I checked what $0.14 and $0.28 per million tokens really buy.

On this page
  1. What exactly shipped on 31 July?
  2. How close is it to Opus 4.8?
  3. Which Opus 4.8 score should you believe?
  4. What do the first independent numbers show?
  5. What does $0.28 per million tokens buy in practice?
  6. The caveats that ship with the discount

DeepSeek shipped the official build of V4 Flash on 31 July 2026 [1], and its launch chart does something unusual: it compares the model with Claude Opus 4.8 on nine benchmarks and loses every row [2]. The excitement comes from the price anyway: $0.14 per million input tokens and $0.28 per million output [3], against $5 and $25 for Opus 4.8 [4].

The model is three days old as I write this, so nothing here is a hands-on verdict. Instead I went through DeepSeek’s launch material, Anthropic’s own numbers for Opus 4.8, the public leaderboards and the first independent measurements, looking for the answer to two questions. How close is “close to Opus 4.8” once you read the footnotes? And what does a discount of 97 to 99 percent actually cost you?

What exactly shipped on 31 July?

DeepSeek-V4-Flash-0731 is the official version of a model that had been running in public preview since 24 April 2026 [1]. DeepSeek states that the new build keeps the architecture and size of the preview and was only re-post-trained, which means the base model is unchanged and only the final training stages were redone [1]. Underneath, V4 Flash is a mixture-of-experts model with 284 billion total parameters, of which 13 billion are active per token, and it accepts a context of one million tokens [5].

The weights are on Hugging Face under an MIT license, so commercial use and self-hosting are unrestricted [2]. On the API side, DeepSeek swapped the build in place: the model string is still deepseek-v4-flash, and existing integrations moved to the new version without a config change [1]. That is convenient for users and awkward for anyone citing results, because “V4 Flash” now names two different builds depending on the date a number was measured. The model exposes a reasoning effort parameter with three levels, low, high and max, instead of shipping a separate reasoning variant [2].

Then there is the price list, which is the reason this release made noise well beyond the model-watcher crowd. Nikkei Asia filed it under an escalating price war in AI models, and reported that DeepSeek plans to introduce peak-hour pricing later [6]. DeepSeek’s own pricing page confirms that plan: during peak hours, prices will double once the policy takes effect, which has not happened yet [3]. The larger sibling, V4 Pro, remains in preview at $0.435 in and $0.87 out, with its official release announced as coming soon [1] [3].

input price, cache miss
0.14 $/MTok
Opus 4.8: $5
output price
0.28 $/MTok
Opus 4.8: $25
context window
1M
max output 384K tokens
active parameters
13B
of 284B total, MIT license

Per million output tokens, Opus 4.8 costs 89 times more than V4 Flash. Whether that ratio matters depends entirely on how much capability sits on the other side of it, which is what the launch table is supposed to establish.

How close is it to Opus 4.8?

Not as close as the summaries floating around suggest. The model card on Hugging Face compares V4-Flash-0731 with Claude Opus 4.8 on nine agent and coding benchmarks, and Opus 4.8 leads every row [2]. Two of the gaps are genuinely small: half a point on Agents’ Last Exam and 2.3 points on Terminal-Bench 2.1. Others are wide: 15.5 points on NL2Repo and 12.1 on DSBench-Hard.

Benchmark V4 Flash 0731Opus 4.8GLM-5.2
Terminal-Bench 2.1 82.7 85.0 (Best value in this row) 81.0
NL2Repo 54.2 69.7 (Best value in this row) 48.9
Cybergym 76.7 83.1 (Best value in this row) -
DeepSWE 54.4 58.0 (Best value in this row) 46.2
Toolathlon Verified 70.3 76.2 (Best value in this row) 59.9
Agents' Last Exam 25.2 25.7 (Best value in this row) 23.8
AutomationBench Public 25.1 27.2 (Best value in this row) 12.9
DSBench FullStack 68.7 71.6 (Best value in this row) 61.8
DSBench Hard 59.6 71.7 (Best value in this row) 54.5
Figure 1. DeepSeek's launch comparison for V4-Flash-0731. All numbers are DeepSeek-run on its own harness at max reasoning effort.

The table has a consistent pattern: V4 Flash sits near Opus 4.8 on terminal-style task benchmarks and falls clearly behind wherever the job involves understanding or building out a repository. Against GLM-5.2, the other Chinese model in the chart, V4 Flash leads all eight rows they share.

To be fair to DeepSeek, its own framing is more modest than the coverage. The card’s headline claim is that the new build outperforms the larger V4 Pro preview despite activating far fewer parameters [2], and printing a comparison you lose on every row is more honest than the usual launch-chart practice of only showing the wins.

The choice of opponent has a quiet problem, though. Opus 4.8 was Anthropic’s best available model when it launched on 28 May 2026 [7], but ten weeks later it sits in the legacy section of Anthropic’s model list [8]. Opus 5 arrived on 24 July at the same $5 and $25, and Anthropic now points new work at it, with Fable 5 above both [9]. Getting near Opus 4.8 in August 2026 means getting near the frontier of May, and the same money now buys a stronger Claude.

Which Opus 4.8 score should you believe?

That depends on who ran the benchmark, and the spread is wide enough to swallow the whole comparison. DeepSeek’s harness gives Opus 4.8 a score of 85.0 on Terminal-Bench 2.1 [2]. The public Terminal-Bench leaderboard lists Opus 4.8 running in Claude Code at 78.9 percent [10]. Anthropic’s own system card reports 74.6 percent, measured at high effort on the Terminus-2 harness [7].

Claude Opus 4.8 on Terminal-Bench 2.1, three measurements DeepSeek's harness reports 85.0, the public leaderboard 78.9, and Anthropic's system card 74.6 for the same model on the same benchmark. DeepSeek Harness, max effort 85.0% Terminal-Bench leaderboard, Claude Code 78.9% Anthropic system card, Terminus-2, high effort 74.6%
Show the data as a table
Source Value
DeepSeek Harness, max effort 85.0%
Terminal-Bench leaderboard, Claude Code 78.9%
Anthropic system card, Terminus-2, high effort 74.6%
Figure 2. One model, one benchmark, three measurements. Terminal-Bench 2.1 scores for Claude Opus 4.8 as reported by three different sources.

That is a spread of 10.4 points for one model on one benchmark, and V4 Flash’s own 82.7 lands inside it. Depending on which Opus number you put next to it, DeepSeek’s model is 2.3 points behind, 3.8 ahead, or 8.1 ahead. A comparison that can produce all three conclusions at once is not really a comparison.

The spread is not evidence of cheating. Each source ran a different setup: a different agent harness, a different effort setting, different timeouts. Anthropic’s system card even notes that Terminal-Bench penalizes slow endpoints, because tasks run against fixed wall-clock limits [7]. Notice that DeepSeek’s harness gave Opus 4.8 a higher score than Anthropic’s own run did, so DeepSeek did not lowball its competitor. The numbers simply come from setups that cannot be compared directly.

What would settle this is a third party running both models through one harness. DeepSeek’s numbers came from its own DeepSeek Harness, which I could not find a public release of, and as of 3 August 2026 I found no independent reproduction of any of the nine rows. Until one exists, the launch table is a claim, not a measurement.

What do the first independent numbers show?

One standardized third-party evaluation exists so far. Artificial Analysis measured V4-Flash-0731 at 50 on its Intelligence Index on 31 July, ten points above the April preview and one point behind GPT-5.6 Luna at max effort [11]. On the same leaderboard as of 3 August, Gemini 3.6 Flash also sits at 50, GLM-5.2 at 51, GPT-5.6 Terra at 55 and Kimi K3 at 57 [12]. The independent placement is therefore the top of the budget class rather than the frontier, and since Opus 4.8 does not appear in that index, no independent number for the Opus comparison exists at all.

The cost side is where the independent data gets interesting. Running the full index cost Artificial Analysis $72.02 with V4 Flash [13], against $176.34 with DeepSeek’s own larger V4 Pro, which scored six points lower at 44 [14]. The small model beating its big sibling was DeepSeek’s central launch claim, and this confirms it on an independent harness: better results at 40 percent of the cost.

Token usage deserves its own sentence, because the cheap rate hides it. V4 Flash needed roughly 206 million output tokens to complete the index, 12 percent fewer than the April build [11], but still more than the 180 million that V4 Pro used [14]. The model is cheap per token, not economical with tokens, so budget estimates based on another model’s token counts will run low.

Two more independent readings fill out the picture. On GDPval-AA v2, a work-task comparison scored by preference votes, the new build reached an Elo of 1559, up from 1189 for the April preview [11], and its hallucination rate on the AA-Omniscience test improved by 12 points to 84 percent, which is progress from a bad starting point rather than a good result [11]. On Arena’s WebDev leaderboard, where humans vote on generated web apps, V4 Flash at high effort held seventh place with an Elo of 1577 on 3 August, against 1705 for Opus 5 Max at the top [15]. And one number is simply missing: no independent speed measurement of the 0731 build existed when I checked, so treat any tokens-per-second figure in circulation as belonging to an older build [16].

What does $0.28 per million tokens buy in practice?

The list price is only the start of the cost model. Cached input costs $0.0028 per million tokens, one fiftieth of the cache-miss rate [3], and agent workflows resend mostly unchanged context on every turn, so the cache rate dominates real bills. In the other direction, the announced peak-hour policy will double prices during Beijing business hours once it activates [3]. Resellers complicate the picture further: at launch, OpenRouter listed V4 Flash at a 38 percent discount, roughly $0.09 in and $0.17 out [17]. If you have seen $0.17 or $0.18 quoted as the model’s output price, that is the reseller promotion, not DeepSeek’s list.

One caveat applies to every cross-vendor token price in this article. Anthropic’s pricing documentation states that Claude models from 4.7 onward use a tokenizer that produces roughly 30 percent more tokens for the same text [4], so per-token ratios between vendors are approximations. A factor of 89 on output survives a 30 percent correction with room to spare.

The early real-world reports match the pricing math. On Hacker News, one engineer posted their 30-day dashboard: $4.55 for 3,467 API requests covering 323 million tokens, with the verdict that the model “has not disappointed at all for coding or review tasks. For everything else, use another model” [18]. That averages about $0.014 per million tokens, below every listed rate except the cache hit, so most of that traffic must have been cached input. The same user describes the working pattern that makes this sustainable: the model is weak at prose, improves noticeably with precise instructions and self-review tooling, and “you must still review 100% of the code as if it’s trying to sell you insurance” [19].

The friction reports are consistent too. Another commenter hit a tight output-token limit through OpenRouter, where a model that talks itself into a long reasoning detour runs out of room before finishing, and still concluded that it “replaced the Kimi models for me” [20]. Simon Willison’s same-day test points at the setting behind that behavior: at the default reasoning effort his result was clearly worse than at high [21]. The practical configuration that emerges from three days of reports: run it at high effort, keep the tasks well-scoped, and keep a stronger model or a human in the review seat.

The caveats that ship with the discount

The benchmark caveat runs through everything above: all nine rows of the launch table are vendor-run on an unpublished harness, and two of the nine benchmarks are DeepSeek’s own internal test sets. That is a reason to wait for independent agent evaluations, not a reason to dismiss the model.

A different kind of caveat is measured rather than suspected. CTGT, a research firm that tests model behavior, found that V4 Flash scored 45.45 points more censored on China-sensitive questions than on structurally identical control questions, testing self-hosted weights to rule out filtering by the API provider [22]. The same study found that a model distilled from V4 Flash outputs showed no significant trace of that behavior [22]. If your product answers questions anywhere near politically sensitive territory, that measurement belongs in your evaluation, and it is a property of the weights, not of DeepSeek’s hosting.

Two practical limits round out the list. The model takes text in and produces text out, with no image input [11], which rules out the screenshot-driven frontend loops that multimodal agent setups lean on. And the official API is only one way to run it: the MIT weights [2] mean OpenRouter and other Western providers already serve the model [17], so DeepSeek’s endpoint, its data jurisdiction and its coming peak-hour pricing are all optional rather than part of the deal.

The genuinely new information from 31 July is not that V4 Flash beats Claude; DeepSeek’s own table says it does not. It is the price at which the model loses. Paying about one percent of Opus rates for a model in this class changes which tasks are worth automating at all, even before the benchmark disputes get settled. My next step is to point it at my usual test repository with reasoning effort at high and a stronger model in the review seat, and to revisit the vendor chart if an independent harness ever runs all of these models side by side.

Sources

  1. Change logDeepSeek API Docs · 2026-07-31
  2. DeepSeek-V4-Flash-0731 model cardDeepSeek on Hugging Face · 2026-07-31
  3. Models & PricingDeepSeek API Docs
  4. PricingClaude Platform Docs
  5. DeepSeek-V4 model cardDeepSeek on Hugging Face
  6. DeepSeek releases beta version of V4 models as AI price war heats upNikkei Asia · 2026-07-31
  7. System Card: Claude Opus 4.8Anthropic · 2026-05-28
  8. Models overviewClaude Platform Docs
  9. Introducing Claude Opus 5Anthropic · 2026-07-24
  10. Terminal-Bench 2.1 leaderboardTerminal-Bench
  11. DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence IndexArtificial Analysis · 2026-07-31
  12. Models leaderboardArtificial Analysis
  13. DeepSeek V4 Flash 0731: intelligence, performance and priceArtificial Analysis
  14. DeepSeek V4 Pro: intelligence, performance and priceArtificial Analysis
  15. Arena leaderboardArena
  16. DeepSeek V4 Flash providersArtificial Analysis
  17. DeepSeek V4 FlashOpenRouter
  18. Comment with 30-day usage figures on the V4 Flash release threadHacker News · 2026-07-31
  19. Comment on working practices with V4 FlashHacker News · 2026-07-31
  20. Comment on output limitsHacker News · 2026-07-31
  21. deepseek-ai/DeepSeek-V4-Flash-0731Simon Willison · 2026-07-31
  22. Distillation censorship transferCTGT · 2026-07-29