Login
Sign Up
Woofun AI reports that DeepSeek released internal performance metrics for agents based on V4-Flash-Vision-Exp, achieving a 2:2 split against Opus 4.8 across four multimodal evaluations. The model scored 27.3 versus 25.7 in Agents' Last Exam and 35.0 versus 34.0 in ZeroBench, while trailing in ApexBench (36.5 vs 39.4) and Chartography (64.3 vs 65.0).
Compared to the text-only V4-Flash-0731, the visual variant improved significantly in multimodal tasks, such as rising from 26.2 to 36.5 in ApexBench. Text-based agent capabilities remained robust, outperforming the previous version in six of seven tasks, including DeepSWE (59.3 vs 54.4) and Toolathlon (75.9). These figures derive from DeepSeek's internal testing using the simplest mode of the DeepSeek Harness.