Skip to content
DigitalNeuron
Modelos e investigación

DeepSeek releases V4.1-Flash with native visual understanding

DeepSeek launched V4.1-Flash on its API, introduced lower prices and outlined the retirement and rerouting of earlier V4 models.

Por DigitalNeuron Desk2 min de lectura

Respuesta rápida

What did DeepSeek announce with the release of DeepSeek-V4.1-Flash?

DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native visual understanding. The company says its asymmetric architecture activates 8 billion parameters for input and 16 billion for output. The model is live through the DeepSeek API, where it carries new peak and off-peak pricing.

Claves

  • DeepSeek says V4.1-Flash is the smallest model in its new architecture family and includes native visual understanding.
  • The 552-billion-parameter mixture-of-experts model activates 8 billion parameters for input and 16 billion for output, according to DeepSeek.
  • DeepSeek says the model’s KV cache requires one-quarter of the HBM and one-eighth of the SSD storage used by the previous generation.
  • V4.1-Flash is available through the DeepSeek API under the model name deepseek-flash.
  • DeepSeek afirmó que todas las solicitudes de deepseek-v4-pro comenzarán a dirigirse a V4.1-Flash a las 04:00 UTC del 14 de septiembre de 2026.

DeepSeek has released DeepSeek-V4.1-Flash, a new mixture-of-experts model with native visual understanding, through its API. The company describes it as the smallest model in a new architecture family designed for greater capability, faster inference, higher throughput and expansion to larger models.

Developers can select the model using deepseek-flash. DeepSeek also published links to the model and its technical report on Hugging Face.

Asymmetric architecture

DeepSeek says V4.1-Flash has 552 billion parameters and uses a new Causal Encoder–Decoder architecture. According to the company, the model activates 8 billion parameters when processing input and 16 billion when producing output.

The company says it applied new pre-training methods and larger-scale reinforcement-learning post-training. DeepSeek claims the resulting benchmark performance exceeds that of flagship models including its own DeepSeek-V4-Pro.

DeepSeek also reported a reduction in the storage required for the model’s key-value cache. Compared with the previous generation, the company says V4.1-Flash needs one-quarter as much high-bandwidth memory and one-eighth as much SSD storage for its KV cache. DeepSeek says this compression can significantly reduce cache-hit costs, which it said often represent a large portion of agent costs.

API changes and model routing

V4.1-Flash is available on the DeepSeek API with native multimodal support. DeepSeek has retired V4-Flash and V4-Flash-Vision-Exp, though requests using deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily route to the new model for compatibility.

The company said tests conducted by multiple parties placed V4.1-Flash ahead of V4-Pro in performance, cost, speed and total runtime. DeepSeek is phasing out V4-Pro.

Starting at 04:00 UTC on September 14, 2026, all requests sent to deepseek-v4-pro will route to V4.1-Flash and use V4.1-Flash rates. DeepSeek said that arrangement will remain in place until V4.1-Pro launches.

DeepSeek identified WorkBuddy, including CodeBuddy, and OpenCode as official partners that fully support V4.1-Flash.

Pricing and deployment

DeepSeek introduced new API pricing alongside the release. The company said peak and off-peak pricing will remain in use to balance demand, with off-peak rates set at 50% of peak rates. The new prices took effect at 04:00 UTC on September 10, 2026.

The company said V4.1-Flash’s architecture allows it to serve more users at a lower cost and that it is passing those savings to API customers.

DeepSeek also said it will work with the open-source community to support V4.1-Flash inference and explore additional deployment options. Its announcement invited organizations planning deployments involving at least 2,000 GPUs and a storage cluster to contact the company.

Source: DeepSeek’s “DeepSeek-V4.1-Flash Release” announcement, published September 10, 2026.

Preguntas frecuentes

What is DeepSeek-V4.1-Flash?
DeepSeek describe V4.1-Flash como el modelo más pequeño de su nueva familia de arquitecturas. Es un modelo de mezcla de expertos con 552.000 millones de parámetros y comprensión visual nativa.
How can developers access V4.1-Flash?
El modelo ya está disponible en la API de DeepSeek con el nombre deepseek-flash. DeepSeek también publicó en Hugging Face enlaces al modelo y a su informe técnico.
What happens to DeepSeek’s earlier V4 models?
DeepSeek retiró V4-Flash y V4-Flash-Vision-Exp, y sus nombres de modelo actuales redirigirán temporalmente a V4.1-Flash. La empresa también indicó que las solicitudes a deepseek-v4-pro se dirigirán a V4.1-Flash desde el 14 de septiembre hasta el lanzamiento de V4.1-Pro.
How does V4.1-Flash pricing work?
DeepSeek afirmó que mantendrá las tarifas de hora punta y valle, y que las tarifas valle se fijarán en el 50 % de las tarifas de hora punta. Los nuevos precios entraron en vigor a las 04:00 UTC del 10 de septiembre de 2026.

Fuentes

  1. DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient | DeepSeek API DocsDeepSeek
Etiquetasdeepseekdeepseek-v4-1-flashmultimodalapimixture-of-expertsopen-source

Lecturas relacionadas

DeepSeek lanza el modelo multimodal V4-Flash-Vision-Exp

DeepSeek lanzó DeepSeek-V4-Flash-Vision-Exp, un modelo multimodal experimental disponible en su plataforma API. La empresa afirma que iguala a DeepSeek-V4-Flash en tareas de texto, al tiempo que incorpora capacidades de visión, y asegura que ofrece una importante mejora en el rendimiento de las pruebas de referencia para agentes multimodales respecto a V4-Flash, acercándose a Opus-4.8. DeepSeek también lanzó una API de archivos gratuita y DeepSeek Harness 0.1.1.

2 min de lectura

Upstage presenta Solar Mini, un modelo de lenguaje compacto y de código abierto

Upstage presentó Solar Mini, un modelo de lenguaje de gran tamaño compacto basado en una arquitectura Llama 2 de 32 capas e inicializado con los pesos de Mistral 7B. La empresa afirma que su método de ampliación de profundidad combina el escalado en profundidad con un preentrenamiento continuo. Solar Mini está disponible públicamente bajo la licencia Apache 2.0.

2 min de lectura

Upstage lanza Solar Pro 4 para tareas de agentes de varios pasos

Upstage lanzó Solar Pro 4, un modelo accesible mediante API y diseñado para tareas de agentes de varios pasos que implican documentos, herramientas y operaciones en terminal. La empresa afirma que admite una ventana de contexto de 512K tokens, genera hasta 128K tokens de salida, trabaja en inglés, coreano y japonés, e informa cuando las pruebas aportadas no permiten sustentar una respuesta.

3 min de lectura