← In the News

DeepSeek will route V4-Pro traffic to its cheaper V4.1-Flash model

Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. · DeepSeek · deepseek.com, September 10, 2026

Machine-readable Download Markdown

DeepSeek released V4.1-Flash, a 552-billion-parameter model built on a new encoder-decoder architecture. DeepSeek says its key-value cache needs a quarter of the HBM and an eighth of the SSD storage used by the previous generation. New pricing for the flash tier took effect at 04:00 UTC today, with off-peak rates set at half of peak.

Starting at 04:00 UTC on September 14, every request still aimed at V4-Pro will route automatically to V4.1-Flash at V4.1-Flash's rates. DeepSeek says it is phasing V4-Pro out entirely. Coding-agent partners WorkBuddy, including CodeBuddy, and OpenCode already support the new model.

"Cache-hit charges often account for a large share of agent costs," DeepSeek wrote. The company presents the smaller cache as a direct cut to the cost of running an agent against its API.

Why it matters: Teams running coding agents against DeepSeek's API on V4-Pro have four days to test the replacement before their traffic moves to a different model at different rates. DeepSeek's cache-cost statement is a vendor claim. Operators should compare it with their agents' actual cache-hit patterns before counting on lower costs.