What is on offer
DeepSeek has opened a limited beta of V4.1 Flash, which it describes as an interim model built on a new architecture with native support for multimodal input, according to TechNode. Developers reach it by switching the model name to deepseek-v4.1-flash-expires-on-0910 — an identifier that states its own expiry. The endpoint goes offline on 10 September.
Beta pricing matches the existing V4 Flash rates, and each account is limited to 20 concurrent requests. DeepSeek’s own claim for the model is the usual trio: stronger performance, faster generation, lower cost.
The architecture claim
This is the part worth attention. DeepSeek says V4.1 Flash is the largest architecture-level change to the V4 line since that series launched in April 2026, and that text, image and speech are handled natively in the same weights rather than through a bolted-on vision component. Its previous multimodal offering, V4-Flash-Vision-Exp, released in August, attached vision to the Flash line externally.

The company has not published an architecture description, a technical report, or benchmark results. A two-day window with a 20-request concurrency cap is not enough for independent evaluation, and DeepSeek has not said it intends to allow one before the model goes generally available.
The speed numbers, with a caveat
Reported throughput figures vary widely and none of them come from DeepSeek. Outlets covering the beta have cited developer measurements ranging from roughly 300 tokens per second to peaks above 500, taken on different prompts under different load. These are not comparable to one another and should not be read as a specification.

Why run a beta like this
A 48-hour public endpoint is a load test with an audience. It gives DeepSeek real traffic patterns on a new architecture without committing to availability, pricing or support, and it generates exactly the developer chatter that a quiet internal test would not. The cost is that nothing verifiable comes out of it.
Where it sits in the line
DeepSeek’s V4 series has been split between Flash, the cheap and fast tier, and Pro, the reasoning tier that went generally available in August with three thinking-effort settings and peak and off-peak pricing. Putting a new architecture into the Flash slot first is the low-risk order: Flash carries the volume traffic, so it surfaces throughput and stability problems quickly, and a Flash model that disappoints costs less reputationally than a Pro one.
It also means the interesting question — whether the architecture holds up on long reasoning traces rather than short completions — is not the one this beta answers.
What to watch
The concrete thing is what replaces the endpoint after 10 September: whether V4.1 Flash returns as a named production model with published pricing and a technical report, or whether the architecture change lands quietly inside an existing model name. The second would be consistent with how DeepSeek has shipped before.