DeepSeek releases V4.1-Flash with native visual understanding
DeepSeek launched V4.1-Flash on its API, introduced lower prices and outlined the retirement and rerouting of earlier V4 models.
Quick answer
What did DeepSeek announce with the release of DeepSeek-V4.1-Flash?
DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native visual understanding. The company says its asymmetric architecture activates 8 billion parameters for input and 16 billion for output. The model is live through the DeepSeek API, where it carries new peak and off-peak pricing.
Key takeaways
- DeepSeek says V4.1-Flash is the smallest model in its new architecture family and includes native visual understanding.
- The 552-billion-parameter mixture-of-experts model activates 8 billion parameters for input and 16 billion for output, according to DeepSeek.
- DeepSeek says the model’s KV cache requires one-quarter of the HBM and one-eighth of the SSD storage used by the previous generation.
- V4.1-Flash is available through the DeepSeek API under the model name deepseek-flash.
- DeepSeek said all deepseek-v4-pro requests will begin routing to V4.1-Flash at 04:00 UTC on September 14, 2026.
DeepSeek has released DeepSeek-V4.1-Flash, a new mixture-of-experts model with native visual understanding, through its API. The company describes it as the smallest model in a new architecture family designed for greater capability, faster inference, higher throughput and expansion to larger models.
Developers can select the model using deepseek-flash. DeepSeek also published links to the model and its technical report on Hugging Face.
Asymmetric architecture
DeepSeek says V4.1-Flash has 552 billion parameters and uses a new Causal Encoder–Decoder architecture. According to the company, the model activates 8 billion parameters when processing input and 16 billion when producing output.
The company says it applied new pre-training methods and larger-scale reinforcement-learning post-training. DeepSeek claims the resulting benchmark performance exceeds that of flagship models including its own DeepSeek-V4-Pro.
DeepSeek also reported a reduction in the storage required for the model’s key-value cache. Compared with the previous generation, the company says V4.1-Flash needs one-quarter as much high-bandwidth memory and one-eighth as much SSD storage for its KV cache. DeepSeek says this compression can significantly reduce cache-hit costs, which it said often represent a large portion of agent costs.
API changes and model routing
V4.1-Flash is available on the DeepSeek API with native multimodal support. DeepSeek has retired V4-Flash and V4-Flash-Vision-Exp, though requests using deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily route to the new model for compatibility.
The company said tests conducted by multiple parties placed V4.1-Flash ahead of V4-Pro in performance, cost, speed and total runtime. DeepSeek is phasing out V4-Pro.
Starting at 04:00 UTC on September 14, 2026, all requests sent to deepseek-v4-pro will route to V4.1-Flash and use V4.1-Flash rates. DeepSeek said that arrangement will remain in place until V4.1-Pro launches.
DeepSeek identified WorkBuddy, including CodeBuddy, and OpenCode as official partners that fully support V4.1-Flash.
Pricing and deployment
DeepSeek introduced new API pricing alongside the release. The company said peak and off-peak pricing will remain in use to balance demand, with off-peak rates set at 50% of peak rates. The new prices took effect at 04:00 UTC on September 10, 2026.
The company said V4.1-Flash’s architecture allows it to serve more users at a lower cost and that it is passing those savings to API customers.
DeepSeek also said it will work with the open-source community to support V4.1-Flash inference and explore additional deployment options. Its announcement invited organizations planning deployments involving at least 2,000 GPUs and a storage cluster to contact the company.
Source: DeepSeek’s “DeepSeek-V4.1-Flash Release” announcement, published September 10, 2026.
Frequently asked questions
- What is DeepSeek-V4.1-Flash?
- DeepSeek describes V4.1-Flash as the smallest model in its new architecture family. It is a 552-billion-parameter mixture-of-experts model with native visual understanding.
- How can developers access V4.1-Flash?
- The model is live on the DeepSeek API under the name deepseek-flash. DeepSeek also linked to the model and its technical report on Hugging Face.
- What happens to DeepSeek’s earlier V4 models?
- DeepSeek retired V4-Flash and V4-Flash-Vision-Exp, with their existing model names temporarily routing to V4.1-Flash. The company also said deepseek-v4-pro requests will route to V4.1-Flash from September 14 until V4.1-Pro launches.
- How does V4.1-Flash pricing work?
- DeepSeek said peak and off-peak pricing will continue, with off-peak rates set at 50% of peak rates. The new pricing took effect at 04:00 UTC on September 10, 2026.