Skip to content
Agents · Aug 1, 2026

DeepSeek V4-Flash 0731 update improves agentic performance without changing architecture or size

Post-training update to DeepSeek V4-Flash raises Terminal-Bench score by 25.8 points and achieves cost-per-task gains via API cache discounts, with open-weights release under MIT license.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • DeepSeek V4-Flash 0731 is a post-training update that improves agentic capabilities without changing model size or architecture.
  • Terminal-Bench score increased by 25.8 points to 82.7, while cost per task remained marginally higher than the prior version.
  • The model is 284B total / 13B active parameters, text-only, with 1M context and $0.14/$0.28 per 1M tokens; API offers a 98% cache-hit discount.
  • Open-weights were released under MIT on Hugging Face with serving details including 256 routed experts and 6 active per token.
  • Community reports note practical integration into coding stacks and renewed relevance for open models amid price compression.

DeepSeek released V4-Flash 0731 as a post-training update that improves agentic capabilities without altering the model’s architecture or size, according to community analyses and DeepSeek’s own statements. Observers highlighted a Terminal-Bench score increase of 25.8 points to 82.7, up from the April preview’s 56.9, as evidence of the upgrade’s impact on agentic tasks.

The model’s specifications remain unchanged: 284B total parameters with 13B active, text-only, and supporting a 1M context window. Pricing on DeepSeek’s API is listed at $0.14 per 1M input tokens and $0.28 per 1M output tokens, with an unusually aggressive 98% cache-hit discount reducing cached token costs to $0.0028 per 1M tokens.

Community evaluators reported additional agentic gains, including a jump in GDPval-AA v2 Elo from 1189 to 1559, Terminal-Bench 2.1 to 79%, a τ³-Bench Banking improvement of 8 points, and a 12% reduction in output-token usage compared to the predecessor. Multiple contributors characterized the gains as a post-training win rather than a scaling-law or pretraining story.

DeepSeek also announced the public-beta launch of the DeepSeek-V4-Flash API, stating that its upgraded agent capabilities surpass V4-Pro-Preview and that the API supports the Responses API format and is "fully adapted for Codex." The company later clarified that these improvements apply only to the Flash API, with V4-Pro API, app, and web remaining unchanged pending an official release.

Open-weights were released under the MIT license on Hugging Face and quickly amplified by community figures. Serving details included 256 routed experts with 6 active per token, a 1M context window, three reasoning-effort levels, and an included DSpark speculative decoding module enabled via a single flag.

Local and quantized deployments followed shortly after. UnslothAI published runnable quantized versions requiring approximately 168GB RAM for lossless 4-bit and 110GB for 3-bit, while additional quantized versions were later shared by other contributors.

Practitioners noted that the cost-performance delta is now large enough that routing and harness choices materially affect engineering workflows. Developers integrated DeepSeek V4-Flash into existing coding stacks, including use within Codex via a router preserving access to multiple models, addition to Hermes Agent, and free public endpoints.

The release renewed discussions about the role of open models in safety and cybersecurity contexts, with some arguing that open models provide defensive advantages and others advocating staged access expansion rather than treating open weights and safety as mutually exclusive.

Sources
  1. 01Latent Space — swyx[AINews] not much happened today
Also on Agents

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.