DeepSeek-V4-Pro launches today
DeepSeek released an experimental multimodal model on Friday and said its agent performance is close to Anthropic’s Opus-4.8. On DeepSeek’s own published table, it beats that model on three benchmarks out of eleven.
DeepSeek-V4-Flash-Vision-Exp is live on the company’s API platform. It takes the text-only V4-Flash and adds the ability to read images and screenshots. It can then act on what it sees. The Hangzhou company also shipped version 0.1.1 of its agent harness, with support built in.
“This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities, including agents, reasoning, and world knowledge,” DeepSeek wrote on X. On multimodal agent benchmarks it “makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8”.
Which Opus, and why it matters less than it looks
Opus-4.8 arrived in May. Anthropic introduced Claude Opus 5 on 24 July, and TNW reported then that Opus 5 matched Fable on coding for half the cost.
Opus-4.8 has not gone anywhere. Anthropic’s model deprecation page lists it as Active, a status the company defines as fully supported and recommended for use, with no retirement earlier than May 2027. Five Opus models carry that status at once.
So DeepSeek picked a current, supported competitor rather than an abandoned one. What the table does not carry is an Opus 5 column, and nobody else has published that comparison either. How V4-Flash-Vision-Exp measures against Anthropic’s newest Opus is simply unknown, and the release does not claim otherwise.
Bloomberg reported the release on Friday, describing Opus 4.8 as an advanced Anthropic model.
What the numbers actually say
DeepSeek published eleven benchmark results. Its new model beats Opus-4.8 on three of them.
It wins DeepSWE by 1.3 points, Agents’ Last Exam by 1.6, and ZeroBench by 1.0. On the other eight it trails. On two of those the margin is wide. NL2Repo has DeepSeek at 57.7 against 69.7, a gap of 12 points. DSBench-Hard has it at 63.6 against 71.7.
The close results are genuinely close. Toolathlon-Verified splits 75.9 to 76.2. Chartography splits 64.3 to 65.0. Terminal Bench 2.1 is 83.9 against 85.0.
One number is worth noting for what it says about the field rather than the contest. On AutomationBench, all three models score in the mid-twenties: 25.7, 25.1 and 27.2. Whatever agents are now good at, that benchmark is not it.
The multimodal leap has an asterisk, and DeepSeek supplied it
The headline claim is the jump in multimodal agent performance over V4-Flash. On ApexBench that is 36.5 against 26.2, and on Agents’ Last Exam 27.3 against 25.2.
DeepSeek’s own footnote explains part of the gap. In those two evaluations, it says, the text-based V4-Flash “ignores multimodal elements contained therein”. The older model is being scored on tests containing images it cannot see.
That does not make the new model’s scores wrong. It does mean the leap is partly a measurement of what happens when you give a blind model an eye test. DeepSeek disclosed it in the table rather than leaving it to be found, which is more than many labs do.
DeepSeek is underselling one result
The company says the vision model “matches” V4-Flash on text. On its own figures it does better than that.
Across the seven text benchmarks, the vision variant beats the text-only model on six. Toolathlon-Verified improves by 5.6 points, DeepSWE by 4.9, DSBench-Hard by 4.0. Cybergym is the exception. There the vision model scores 75.3 against 76.7, so adding sight cost it 1.4 points on a security benchmark.
Every one of these figures comes from DeepSeek. The company states that it evaluated its own models using its own harness in minimal mode, with temperature at 1.0 and top_p at 0.95. Vendor benchmarks are normal, and the settings are disclosed. They are still the vendor’s.
Why the comparison matters commercially
The reason any of this lands is price. Research this month found V4-Flash to be the cheapest well-known model to run. A million words costs about 87 cents from DeepSeek against roughly $50 from Anthropic, a gap corporate buyers have already noticed.
Against that spread, trailing by a point or two is a commercial argument rather than a defeat. Losing NL2Repo by 12 points is a different matter. Repository-scale work is exactly what enterprises are buying agents to do.
DeepSeek separately confirmed on its own site that the official version of V4-Pro has shipped, with what it calls significantly enhanced agent capabilities, Responses API support and Codex integration. That is the model most enterprise buyers would actually be weighing, and it is not in Friday’s table either.
Read the claim, then read the table
None of this is unusual. Labs pick favourable baselines. The desk has been here recently, when Alibaba called Qwen the world’s most downloaded open model, which was true by a smaller margin than claimed.
DeepSeek’s statement is defensible on its own terms. It said close to Opus-4.8, and on several benchmarks it is close to Opus-4.8. It did not say close to Anthropic’s newest model, and it did not claim to lead on the set.
The test that would settle it has a name and no results. Somebody needs to run V4-Flash-Vision-Exp against Opus 5 on the same harness. Until then the honest summary is narrower than the headlines. An experimental Chinese multimodal model sits within a few points of a supported American model on most of a benchmark set, wins three of eleven, costs a fraction as much, and trails by twelve points on the hardest task in it.
Get the TNW newsletter
Get the most important tech news in your inbox each week.