DeepSeek Releases V4.1 Flash: 552B Multimodal MoE and Aggressive API Price Cuts
DeepSeek released V4.1 Flash, a 552B multimodal MoE activating 8B/16B parameters, alongside major API price reductions and upcoming V4 Pro routing.
On September 10, 2026, Chinese AI research lab DeepSeek officially launched DeepSeek V4.1 Flash, its next-generation native multimodal Mixture-of-Experts (MoE) model, along with a production-ready developer API. Announced via DeepSeek's official news portal and API documentation changelogs, this release combines compute efficiency across massive sparse parameters with aggressive API pricing reductions that significantly disrupt the developer ecosystem.

Image source: DeepSeek official announcement
Designed to process both text and visual inputs natively within a unified pipeline, V4.1 Flash dramatically lowers token inference costs while maintaining high-fidelity multimodal reasoning. Alongside the new model release, DeepSeek also published a comprehensive pricing schedule and confirmed a clear migration roadmap for existing V4 Pro API workloads.
552B MoE Architecture and Causal-Encoder-Decoder Efficiency
The most prominent technical advancement in DeepSeek V4.1 Flash is its native multimodal Mixture-of-Experts architecture operating across a total footprint of 552 billion parameters.
The model utilizes a Causal-Encoder-Decoder design that decouples and independently optimizes prompt ingestion from autoregressive token generation:
- Prompt Input Processing: The model selectively activates only 8 billion parameters out of the 552B total during prompt parsing.
- Response Generation: During text and multimodal token emission, it activates only 16 billion parameters.
This sparse routing strategy ensures that developers benefit from the extensive pretraining knowledge and generalization capabilities of a 552B parameter foundation model, while keeping compute overhead and VRAM bandwidth consumption comparable to much smaller lightweight models. As a result, engineering teams achieve rapid time-to-first-token (TTFT) and high sustained token throughput under production loads.
API Pricing Structure and Peak vs. Off-Peak Rates
The newly established V4.1 Flash pricing took effect immediately on September 10, 2026, at 04:00 UTC (12:00 Beijing Time / CST).
The baseline off-peak pricing structure per 1 million tokens is tiered as follows:
- Cache Hit Input: 0.02 CNY per 1M tokens
- Cache Miss Input: 1.00 CNY per 1M tokens
- Output Generation: 4.00 CNY per 1M tokens
During designated peak hours, standard rates double across all prompt and output tiers. To prevent operational ambiguity across international engineering teams, DeepSeek explicitly specified rollout and transition milestones in both UTC and Beijing Time. Developers running batch evaluation pipelines, fine-tuning evaluations, or large-scale document indexing workflows can achieve significant cost savings by scheduling tasks within off-peak windows.
V4 Pro Routing Transition and Reseller Promotion Context
DeepSeek has also established an orderly deprecation and transition schedule for developers currently building on earlier model variants.
Effective September 14, 2026, at 12:00 Beijing Time (04:00 UTC), all incoming API requests directed to the deepseek-v4-pro endpoint will automatically route to V4.1 Flash. Once this transition takes effect, those workloads will be billed under the new Flash pricing schedule, substantially reducing recurring infrastructure costs for existing enterprise integrations.
Meanwhile, claims circulated by third-party API reseller B.AI (@BAI_AGI) regarding additional 90% discounts and deposit matching bonuses represent independent reseller marketing promotions. Because these promotional incentives are not part of DeepSeek's official release announcement, developers should evaluate official platform documentation independently from third-party reseller campaigns when planning production budgets.
Sources
This article is based on verified technical release notes and API documentation from DeepSeek, alongside relevant developer community announcements.
- DeepSeek Official Announcement: DeepSeek V4.1 Flash Release and Pricing Documentation
- B.AI Community Announcement: B.AI Release Notice and Reseller Promotion on X