DeepSeek just dropped what might be the most cost-efficient frontier AI model on the market. The Chinese research lab’s new V4-Flash can process both text and images, performs competitively with Anthropic’s Claude Opus 4.8 on reasoning and coding benchmarks, and does it all for roughly $0.28 per output. Anthropic charges somewhere in the $25 to $30 range for comparable work.
What DeepSeek actually built
The company released preview versions of DeepSeek-V4-Pro and V4-Flash on April 24, 2026, the latest entries in its V4 model family. The headline feature is multimodal capability, meaning these models can interpret images alongside text prompts.
The technical architecture behind V4-Flash is where things get interesting. The model uses approximately 90 KV cache entries to process images, compared to roughly 870 for Claude models. KV cache is essentially the model’s working memory during inference. Fewer entries means less computational overhead, which translates directly into lower costs and faster processing.
DeepSeek’s models also support long-context handling of up to 1 million tokens. The models are compatible with both OpenAI and Anthropic APIs, making them relatively straightforward drop-in replacements for developers already building on those platforms.
Subsequent updates have followed the initial release, including a V4-Flash-0731 version and an experimental vision model dubbed deepseek-v4-flash-vision-exp, suggesting the lab is iterating quickly on its multimodal capabilities.
The cost gap that keeps widening
DeepSeek has been building toward this moment since January 2025, when its R1 model first rattled the AI industry. That release established a template the company has now refined: match or approach frontier performance while dramatically undercutting on price.
V4-Flash performs on par with Claude Opus 4.8 in benchmarks covering reasoning, coding, and agentic tasks while operating at approximately 99% less cost. Much of this efficiency comes from DeepSeek’s use of Mixture-of-Experts architecture, a design approach where only a subset of the model’s parameters activate for any given task.
A company running thousands of AI inference calls per day could see its compute bill drop by orders of magnitude simply by switching providers. At $0.28 versus $25 to $30 per output, the math doesn’t require a spreadsheet to figure out.
Why this matters beyond the benchmarks
Reports of distillation attacks, where one lab’s model outputs are used to train a competitor’s system, have added tension to the relationship between DeepSeek and Western AI companies. API compatibility with OpenAI and Anthropic means switching costs are relatively low for developers evaluating alternatives.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
18








English (US) ·