August 1, 2026 · Digital Applied / Artificial Analysis
DeepSeek V4 Flash 0731 Official Release: DeepSWE Score Jumps 7x to 54.4 After Re-Post-Training
My take: DeepSeek released the official version of its V4 Flash model yesterday with a technical result worth noting: without changing the architecture or model size, only with an additional post-training cycle, the score on DeepSWE, the most demanding agent coding benchmark, jumped from 7.3 to 54.4. A sevenfold improvement in a single update.
For those building AI workflows or using assisted programming tools, that is significant: cheaper models are reaching capabilities that until recently only the most expensive models had. V4 Flash still costs $0.14 per million input tokens, less than a tenth of what leading OpenAI or Anthropic models charge.
The capability race no longer happens only at the expensive end of the market. If you are choosing which model to use in your projects or tools, checking agent benchmarks every few weeks is already part of the work for those who build with AI seriously.
What capability benchmark do you use to decide which model to integrate into what you build or your workflow?
Want to use these tools? See the unbiased reviews or back to the news.