DeepSeek releases V4.1-Flash with native visual understanding
DeepSeek launched V4.1-Flash on its API, introduced lower prices and outlined the retirement and rerouting of earlier V4 models.
Respuesta rápida
What did DeepSeek announce with the release of DeepSeek-V4.1-Flash?
DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native visual understanding. The company says its asymmetric architecture activates 8 billion parameters for input and 16 billion for output. The model is live through the DeepSeek API, where it carries new peak and off-peak pricing.
Claves
- DeepSeek says V4.1-Flash is the smallest model in its new architecture family and includes native visual understanding.
- The 552-billion-parameter mixture-of-experts model activates 8 billion parameters for input and 16 billion for output, according to DeepSeek.
- DeepSeek says the model’s KV cache requires one-quarter of the HBM and one-eighth of the SSD storage used by the previous generation.
- V4.1-Flash is available through the DeepSeek API under the model name deepseek-flash.
- DeepSeek afirmó que todas las solicitudes de deepseek-v4-pro comenzarán a dirigirse a V4.1-Flash a las 04:00 UTC del 14 de septiembre de 2026.
DeepSeek has released DeepSeek-V4.1-Flash, a new mixture-of-experts model with native visual understanding, through its API. The company describes it as the smallest model in a new architecture family designed for greater capability, faster inference, higher throughput and expansion to larger models.
Developers can select the model using deepseek-flash. DeepSeek also published links to the model and its technical report on Hugging Face.
Asymmetric architecture
DeepSeek says V4.1-Flash has 552 billion parameters and uses a new Causal Encoder–Decoder architecture. According to the company, the model activates 8 billion parameters when processing input and 16 billion when producing output.
The company says it applied new pre-training methods and larger-scale reinforcement-learning post-training. DeepSeek claims the resulting benchmark performance exceeds that of flagship models including its own DeepSeek-V4-Pro.
DeepSeek also reported a reduction in the storage required for the model’s key-value cache. Compared with the previous generation, the company says V4.1-Flash needs one-quarter as much high-bandwidth memory and one-eighth as much SSD storage for its KV cache. DeepSeek says this compression can significantly reduce cache-hit costs, which it said often represent a large portion of agent costs.
API changes and model routing
V4.1-Flash is available on the DeepSeek API with native multimodal support. DeepSeek has retired V4-Flash and V4-Flash-Vision-Exp, though requests using deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily route to the new model for compatibility.
The company said tests conducted by multiple parties placed V4.1-Flash ahead of V4-Pro in performance, cost, speed and total runtime. DeepSeek is phasing out V4-Pro.
Starting at 04:00 UTC on September 14, 2026, all requests sent to deepseek-v4-pro will route to V4.1-Flash and use V4.1-Flash rates. DeepSeek said that arrangement will remain in place until V4.1-Pro launches.
DeepSeek identified WorkBuddy, including CodeBuddy, and OpenCode as official partners that fully support V4.1-Flash.
Pricing and deployment
DeepSeek introduced new API pricing alongside the release. The company said peak and off-peak pricing will remain in use to balance demand, with off-peak rates set at 50% of peak rates. The new prices took effect at 04:00 UTC on September 10, 2026.
The company said V4.1-Flash’s architecture allows it to serve more users at a lower cost and that it is passing those savings to API customers.
DeepSeek also said it will work with the open-source community to support V4.1-Flash inference and explore additional deployment options. Its announcement invited organizations planning deployments involving at least 2,000 GPUs and a storage cluster to contact the company.
Source: DeepSeek’s “DeepSeek-V4.1-Flash Release” announcement, published September 10, 2026.
Preguntas frecuentes
- What is DeepSeek-V4.1-Flash?
- DeepSeek describe V4.1-Flash como el modelo más pequeño de su nueva familia de arquitecturas. Es un modelo de mezcla de expertos con 552.000 millones de parámetros y comprensión visual nativa.
- How can developers access V4.1-Flash?
- El modelo ya está disponible en la API de DeepSeek con el nombre deepseek-flash. DeepSeek también publicó en Hugging Face enlaces al modelo y a su informe técnico.
- What happens to DeepSeek’s earlier V4 models?
- DeepSeek retiró V4-Flash y V4-Flash-Vision-Exp, y sus nombres de modelo actuales redirigirán temporalmente a V4.1-Flash. La empresa también indicó que las solicitudes a deepseek-v4-pro se dirigirán a V4.1-Flash desde el 14 de septiembre hasta el lanzamiento de V4.1-Pro.
- How does V4.1-Flash pricing work?
- DeepSeek afirmó que mantendrá las tarifas de hora punta y valle, y que las tarifas valle se fijarán en el 50 % de las tarifas de hora punta. Los nuevos precios entraron en vigor a las 04:00 UTC del 10 de septiembre de 2026.