DeepSeek Vision Agent Scores Match Opus 4.8 with 2:2 Split
2026-08-21 17:30

Woofun AI reports that DeepSeek released internal performance metrics for agents based on V4-Flash-Vision-Exp, achieving a 2:2 split against Opus 4.8 across four multimodal evaluations. The model scored 27.3 versus 25.7 in Agents' Last Exam and 35.0 versus 34.0 in ZeroBench, while trailing in ApexBench (36.5 vs 39.4) and Chartography (64.3 vs 65.0).

Compared to the text-only V4-Flash-0731, the visual variant improved significantly in multimodal tasks, such as rising from 26.2 to 36.5 in ApexBench. Text-based agent capabilities remained robust, outperforming the previous version in six of seven tasks, including DeepSWE (59.3 vs 54.4) and Toolathlon (75.9). These figures derive from DeepSeek's internal testing using the simplest mode of the DeepSeek Harness.

Disclaimer: Views are the author's own and do not represent the platform. Do not reproduce without permission. Content is for reference only, not investment advice. Trade at your own risk.
Tags:
DeepSeek
V4-Flash-Vision-Exp
Opus 4.8
V4-Flash-0731
DeepSeek Harness
ChainCatcher
Share:
back