A beta that lasted 48 hours
DeepSeek made V4.1 Flash generally available on 10 September, two days after opening a beta endpoint that was scheduled to disappear on the same date. The company had named the beta model deepseek-v4.1-flash-expires-on-0910, so the shutdown date was published before the model was.
In its change log, DeepSeek describes V4.1 Flash as the smallest model in a new architecture family, with native multimodal visual understanding built into the base model rather than bolted on. The company says the design targets a higher capability ceiling, faster inference, higher throughput and the ability to scale to larger models — which is a statement about what comes next as much as about what shipped.
What replaces what
Two model names disappear into the new one. Requests to deepseek-v4-flash and deepseek-v4-flash-vision-exp are now routed to V4.1 Flash. The separate experimental vision endpoint is the clearer signal: image understanding is no longer a side branch of the Flash line.

The larger change is at the top of the range. DeepSeek says that from 14 September, requests to deepseek-v4-pro will also be routed to V4.1 Flash and billed at the Flash price, retiring V4 Pro until a future V4.1 Pro arrives. Developers who chose Pro for its capability will be served a smaller model until that replacement lands, and the company has not said when it will.
Prices and context
DeepSeek’s pricing page lists V4.1 Flash with a one-million-token context window. Off-peak, input costs $0.15 per million tokens on a cache miss and $0.003 on a cache hit, with output at $0.60 per million. Peak rates are double: $0.30, $0.006 and $1.20. DeepSeek defines peak as 01:00–04:00 and 06:00–10:00 UTC on weekdays, and charges off-peak rates the rest of the time.

The company says prices were reduced with the release, without publishing a comparison table. The cache-hit rate is the number that stands out: at $0.003 per million tokens off-peak, repeated prompt prefixes cost almost nothing, which is the shape of pricing aimed at agent loops and long documents rather than one-off chat.
What to watch
DeepSeek has not published benchmark results for V4.1 Flash alongside the release, and any figures circulating from the beta window are user measurements of throughput, not capability scores. The two things worth watching are whether independent evaluations confirm the architecture claim, and when V4.1 Pro appears — because until it does, the company’s most capable public model is a Flash-tier one.