DeepSeek has launched the official public testing API for DeepSeek-V4-Flash on July 31, 2026.
The release reveals enhanced agent capabilities that approach previous preview standards, despite maintaining a compact 13-billion active parameter architecture.
>>> Arsenal Close to Signing Center-Back After Saliba Injury
The updated Flash model achieved a score of 82.7 on Terminal Bench 2.1, alongside 54.2 on NL2Repo, 76.7 on Cybergym, and 70.3 on verified Toolathlon tasks.
In comparison, the earlier V4-Pro preview version scored 67.9 on Terminal Bench 2.0.
Public testing remains restricted to the API, leaving mobile app and web interface integration unavailable. This limitation persists while the DeepSeek-V4-Pro official release is prepared.
Model Consistency and Engineering
DeepSeek stated in an official update log that the model structure and size of DeepSeek-V4-Flash-0731 remain consistent with DeepSeek-V4-Flash-preview, with only post-training being re-conducted.
Engineering efforts behind the update rely on the DeepSeek Harness framework, operated by a dedicated Agent Harness team established in March under lead Cui Tianyi.
>>> West Ham Targets Celtic Duo Engels and Hatate
The official Code Agent benchmark evaluations utilized the minimalist DeepSeek Harness framework configuration running at maximum settings with a temperature parameter of 1.0.
Liang Wenfeng described the progression toward artificial general intelligence, saying, "AI currently does not lack taste and intuition; what it lacks is the ability to learn continuously."
He added, "Investors look at Agents, while we look at how to solve learning."
Market shifts indicate an impending release for the primary V4 Pro model.
Cloud provider Silicon Flow announced price increases for DeepSeek V4 Pro cache hits effective August 3, raising costs from 0.1 yuan to 1.0 yuan per million tokens.
>>> Brighton Pride 2026 Marks 35th Anniversary with Weekend Celebrations
Developer community testing confirmed that select V4 API endpoints successfully processed queries regarding recent unannounced model release dates during limited grayscale deployments.