Skip to content
DigitalNeuron
Models & research

DeepSeek releases V4.1-Flash with native visual understanding

DeepSeek launched V4.1-Flash on its API, introduced lower prices and outlined the retirement and rerouting of earlier V4 models.

By DigitalNeuron Desk2 min read

Quick answer

What did DeepSeek announce with the release of DeepSeek-V4.1-Flash?

DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native visual understanding. The company says its asymmetric architecture activates 8 billion parameters for input and 16 billion for output. The model is live through the DeepSeek API, where it carries new peak and off-peak pricing.

Key takeaways

  • DeepSeek says V4.1-Flash is the smallest model in its new architecture family and includes native visual understanding.
  • The 552-billion-parameter mixture-of-experts model activates 8 billion parameters for input and 16 billion for output, according to DeepSeek.
  • DeepSeek says the model’s KV cache requires one-quarter of the HBM and one-eighth of the SSD storage used by the previous generation.
  • V4.1-Flash is available through the DeepSeek API under the model name deepseek-flash.
  • DeepSeek said all deepseek-v4-pro requests will begin routing to V4.1-Flash at 04:00 UTC on September 14, 2026.

DeepSeek has released DeepSeek-V4.1-Flash, a new mixture-of-experts model with native visual understanding, through its API. The company describes it as the smallest model in a new architecture family designed for greater capability, faster inference, higher throughput and expansion to larger models.

Developers can select the model using deepseek-flash. DeepSeek also published links to the model and its technical report on Hugging Face.

Asymmetric architecture

DeepSeek says V4.1-Flash has 552 billion parameters and uses a new Causal Encoder–Decoder architecture. According to the company, the model activates 8 billion parameters when processing input and 16 billion when producing output.

The company says it applied new pre-training methods and larger-scale reinforcement-learning post-training. DeepSeek claims the resulting benchmark performance exceeds that of flagship models including its own DeepSeek-V4-Pro.

DeepSeek also reported a reduction in the storage required for the model’s key-value cache. Compared with the previous generation, the company says V4.1-Flash needs one-quarter as much high-bandwidth memory and one-eighth as much SSD storage for its KV cache. DeepSeek says this compression can significantly reduce cache-hit costs, which it said often represent a large portion of agent costs.

API changes and model routing

V4.1-Flash is available on the DeepSeek API with native multimodal support. DeepSeek has retired V4-Flash and V4-Flash-Vision-Exp, though requests using deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily route to the new model for compatibility.

The company said tests conducted by multiple parties placed V4.1-Flash ahead of V4-Pro in performance, cost, speed and total runtime. DeepSeek is phasing out V4-Pro.

Starting at 04:00 UTC on September 14, 2026, all requests sent to deepseek-v4-pro will route to V4.1-Flash and use V4.1-Flash rates. DeepSeek said that arrangement will remain in place until V4.1-Pro launches.

DeepSeek identified WorkBuddy, including CodeBuddy, and OpenCode as official partners that fully support V4.1-Flash.

Pricing and deployment

DeepSeek introduced new API pricing alongside the release. The company said peak and off-peak pricing will remain in use to balance demand, with off-peak rates set at 50% of peak rates. The new prices took effect at 04:00 UTC on September 10, 2026.

The company said V4.1-Flash’s architecture allows it to serve more users at a lower cost and that it is passing those savings to API customers.

DeepSeek also said it will work with the open-source community to support V4.1-Flash inference and explore additional deployment options. Its announcement invited organizations planning deployments involving at least 2,000 GPUs and a storage cluster to contact the company.

Source: DeepSeek’s “DeepSeek-V4.1-Flash Release” announcement, published September 10, 2026.

Frequently asked questions

What is DeepSeek-V4.1-Flash?
DeepSeek describes V4.1-Flash as the smallest model in its new architecture family. It is a 552-billion-parameter mixture-of-experts model with native visual understanding.
How can developers access V4.1-Flash?
The model is live on the DeepSeek API under the name deepseek-flash. DeepSeek also linked to the model and its technical report on Hugging Face.
What happens to DeepSeek’s earlier V4 models?
DeepSeek retired V4-Flash and V4-Flash-Vision-Exp, with their existing model names temporarily routing to V4.1-Flash. The company also said deepseek-v4-pro requests will route to V4.1-Flash from September 14 until V4.1-Pro launches.
How does V4.1-Flash pricing work?
DeepSeek said peak and off-peak pricing will continue, with off-peak rates set at 50% of peak rates. The new pricing took effect at 04:00 UTC on September 10, 2026.

Sources

  1. DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient | DeepSeek API DocsDeepSeek
Tagsdeepseekdeepseek-v4-1-flashmultimodalapimixture-of-expertsopen-source

Related reading

DeepSeek Releases V4-Flash-Vision-Exp Multimodal Model

DeepSeek released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model on its API platform. The company says it matches DeepSeek-V4-Flash on text tasks while adding vision, and claims a major jump in multimodal agent benchmark performance over V4-Flash, near Opus-4.8. DeepSeek also launched a free Files API and DeepSeek Harness 0.1.1.

1 min read

Upstage launches Solar Pro 4 for multi-step agent work

Upstage launched Solar Pro 4, an API-accessible model designed for multi-step agent work involving documents, tools and terminal tasks. The company says it supports a 512K-token context window, produces up to 128K output tokens, handles English, Korean and Japanese, and reports when supplied evidence cannot support an answer.

2 min read