Skip to content
DigitalNeuron
모델·연구

DeepSeek releases V4.1-Flash with native visual understanding

DeepSeek launched V4.1-Flash on its API, introduced lower prices and outlined the retirement and rerouting of earlier V4 models.

DigitalNeuron Desk2분 읽기

한 줄 답

What did DeepSeek announce with the release of DeepSeek-V4.1-Flash?

DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native visual understanding. The company says its asymmetric architecture activates 8 billion parameters for input and 16 billion for output. The model is live through the DeepSeek API, where it carries new peak and off-peak pricing.

핵심 요약

  • DeepSeek says V4.1-Flash is the smallest model in its new architecture family and includes native visual understanding.
  • The 552-billion-parameter mixture-of-experts model activates 8 billion parameters for input and 16 billion for output, according to DeepSeek.
  • DeepSeek says the model’s KV cache requires one-quarter of the HBM and one-eighth of the SSD storage used by the previous generation.
  • V4.1-Flash is available through the DeepSeek API under the model name deepseek-flash.
  • DeepSeek said all deepseek-v4-pro requests will begin routing to V4.1-Flash at 04:00 UTC on September 14, 2026.

DeepSeek has released DeepSeek-V4.1-Flash, a new mixture-of-experts model with native visual understanding, through its API. The company describes it as the smallest model in a new architecture family designed for greater capability, faster inference, higher throughput and expansion to larger models.

Developers can select the model using deepseek-flash. DeepSeek also published links to the model and its technical report on Hugging Face.

Asymmetric architecture

DeepSeek says V4.1-Flash has 552 billion parameters and uses a new Causal Encoder–Decoder architecture. According to the company, the model activates 8 billion parameters when processing input and 16 billion when producing output.

The company says it applied new pre-training methods and larger-scale reinforcement-learning post-training. DeepSeek claims the resulting benchmark performance exceeds that of flagship models including its own DeepSeek-V4-Pro.

DeepSeek also reported a reduction in the storage required for the model’s key-value cache. Compared with the previous generation, the company says V4.1-Flash needs one-quarter as much high-bandwidth memory and one-eighth as much SSD storage for its KV cache. DeepSeek says this compression can significantly reduce cache-hit costs, which it said often represent a large portion of agent costs.

API changes and model routing

V4.1-Flash is available on the DeepSeek API with native multimodal support. DeepSeek has retired V4-Flash and V4-Flash-Vision-Exp, though requests using deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily route to the new model for compatibility.

The company said tests conducted by multiple parties placed V4.1-Flash ahead of V4-Pro in performance, cost, speed and total runtime. DeepSeek is phasing out V4-Pro.

Starting at 04:00 UTC on September 14, 2026, all requests sent to deepseek-v4-pro will route to V4.1-Flash and use V4.1-Flash rates. DeepSeek said that arrangement will remain in place until V4.1-Pro launches.

DeepSeek identified WorkBuddy, including CodeBuddy, and OpenCode as official partners that fully support V4.1-Flash.

Pricing and deployment

DeepSeek introduced new API pricing alongside the release. The company said peak and off-peak pricing will remain in use to balance demand, with off-peak rates set at 50% of peak rates. The new prices took effect at 04:00 UTC on September 10, 2026.

The company said V4.1-Flash’s architecture allows it to serve more users at a lower cost and that it is passing those savings to API customers.

DeepSeek also said it will work with the open-source community to support V4.1-Flash inference and explore additional deployment options. Its announcement invited organizations planning deployments involving at least 2,000 GPUs and a storage cluster to contact the company.

Source: DeepSeek’s “DeepSeek-V4.1-Flash Release” announcement, published September 10, 2026.

자주 묻는 질문

What is DeepSeek-V4.1-Flash?
DeepSeek describes V4.1-Flash as the smallest model in its new architecture family. It is a 552-billion-parameter mixture-of-experts model with native visual understanding.
How can developers access V4.1-Flash?
The model is live on the DeepSeek API under the name deepseek-flash. DeepSeek also linked to the model and its technical report on Hugging Face.
What happens to DeepSeek’s earlier V4 models?
DeepSeek retired V4-Flash and V4-Flash-Vision-Exp, with their existing model names temporarily routing to V4.1-Flash. The company also said deepseek-v4-pro requests will route to V4.1-Flash from September 14 until V4.1-Pro launches.
How does V4.1-Flash pricing work?
DeepSeek said peak and off-peak pricing will continue, with off-peak rates set at 50% of peak rates. The new pricing took effect at 04:00 UTC on September 10, 2026.

출처

  1. DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient | DeepSeek API DocsDeepSeek
태그deepseekdeepseek-v4-1-flashmultimodalapimixture-of-expertsopen-source

함께 읽기

DeepSeek, V4-Flash-Vision-Exp 멀티모달 모델 공개

딥시크(DeepSeek)가 자사 API 플랫폼에서 실험적 멀티모달 모델인 DeepSeek-V4-Flash-Vision-Exp를 공개했다. 회사 측은 텍스트 작업에서 DeepSeek-V4-Flash와 동등한 성능을 보이면서 비전 기능을 추가했으며, V4-Flash 대비 멀티모달 에이전트 벤치마크 성능이 크게 향상되어 Opus-4.8에 근접했다고 밝혔다. 딥시크는 또한 무료 Files API와 DeepSeek Harness 0.1.1도 함께 출시했다.

2분 읽기

업스테이지, 소형 오픈소스 언어 모델 ‘Solar Mini’ 공개

업스테이지는 32개 층의 Llama 2 아키텍처를 기반으로 하고 Mistral 7B 가중치로 초기화한 소형 대규모 언어 모델 Solar Mini를 공개했다. 회사 측은 자체 뎁스 업스케일링 기법이 깊이 방향 확장과 지속적 사전학습을 결합한다고 설명했다. Solar Mini는 Apache 2.0 라이선스로 공개돼 있다.

2분 읽기