DeepSeek just dropped a model that processes a million tokens of context while activating fewer parameters than some open-source models released two years ago. The V4.1-Flash, launched on September 10, represents the Chinese AI startup’s latest bid to rewrite the economics of large language models.
The model packs 552 billion parameters into a Mixture-of-Experts (MoE) architecture, but only fires up about 8 billion of them for input tasks and 16 billion for output. The result is a model that punches well above what its active compute footprint would suggest.
The architecture that makes it work
V4.1-Flash introduces what DeepSeek calls an asymmetric Causal Encoder-Decoder architecture, processing input and generating output through different pathways optimized for each task, rather than running everything through a single pipeline.
The context window stretches to 1 million tokens. Supporting that massive context is a KV cache that consumes approximately 890 bytes per token, about one-quarter of what the prior V4-Flash model required.
That cache reduction matters more than it might sound. KV cache is the memory bottleneck that determines how many concurrent users a model can serve and how long their conversations can run. Cutting it by 75% means operators can serve roughly four times as many users on the same hardware, or handle contexts four times as long without upgrading their GPU clusters.
On benchmarks, DeepSeek claims V4.1-Flash outperforms the company’s own V4-Pro model and competes directly with GPT-5.6 and Kimi K3. The company is already routing requests from V4-Pro to the new model, with updated pricing taking effect on September 14.
Open weights, open strategy
DeepSeek released the model’s weights under an MIT license on Hugging Face. The new model is accessible through DeepSeek’s API under the endpoint “deepseek-flash.” The combination of open weights and API access creates a two-track adoption path: developers who want to run inference on their own infrastructure can download and deploy locally, while those who prefer managed services can call the API at DeepSeek’s new pricing tiers.
IPO implications and competitive positioning
The timing of the V4.1-Flash launch is not accidental. DeepSeek is expected to go public on Shanghai’s STAR Market, and demonstrating continued technical momentum is the kind of thing that makes roadshow presentations more convincing.
The multimodal understanding built into V4.1-Flash, spanning both visual and text data, narrows the differentiation opportunities available to rivals including OpenAI and Moonshot AI, which develops Kimi K3.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

41 minutes ago
23









English (US) ·