Skip to content
DigitalNeuron

Models & research

New frontier models, benchmarks, training methods and the research behind them.

14 articles

DeepSeek releases V4.1-Flash with native visual understanding

DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native visual understanding. The company says its asymmetric architecture activates 8 billion parameters for input and 16 billion for output. The model is live through the DeepSeek API, where it carries new peak and off-peak pricing.

2 min read

Upstage launches Solar Pro 4 for multi-step agent work

Upstage launched Solar Pro 4, an API-accessible model designed for multi-step agent work involving documents, tools and terminal tasks. The company says it supports a 512K-token context window, produces up to 128K output tokens, handles English, Korean and Japanese, and reports when supplied evidence cannot support an answer.

2 min read

Fine-tuning, retrieval or a better prompt: how to choose

Diagnose the failure first. If the model does not know something, that is a knowledge gap and retrieval fixes it. If the model knows but answers in the wrong shape, tone or format, that is a behaviour gap and prompting fixes it first, fine-tuning second. If the model fails at a specialised skill after both, fine-tuning is the remaining option. Prompting is cheapest and reversible, retrieval is the right default for facts, and fine-tuning is the most expensive and least reversible of the three.

5 min read

Analysis: long context did not kill retrieval — it changed what retrieval is for

No, but they change its job. Filling a very large window degrades accuracy on information buried in the middle, multiplies latency and cost on every request, and makes it hard to say which source an answer came from. Retrieval remains the right default for large or changing corpora, for anything that needs citations or access control, and for cost-sensitive high-volume paths. Long context is now best used for whole-document reasoning, for agent working memory within a task, and as the second stage after retrieval has narrowed the field.

6 min read

Analysis: public benchmarks stopped predicting production quality — how to build an eval set that does

Public benchmarks measure narrow, static, widely-published tasks, and they are increasingly contaminated by training data and optimised for directly. Production quality depends on your prompts, your documents, your tool schemas and your failure tolerance, none of which any leaderboard measures. The practical answer is a private eval set of 50 to 300 real cases with recorded expected behaviour, run on every model and prompt change, scored by exact checks where possible and by a rubric-driven model judge where not.

6 min read

Anthropic Publishes Second Economic Index Report on Claude 3.7 Sonnet Usage

Anthropic's second Economic Index report found rising Claude.ai usage in coding, education, science and healthcare after the launch of Claude 3.7 Sonnet. Extended thinking mode was used most by technical occupations. Augmentation held steady at 57% of usage, and Anthropic released a new 630-category usage taxonomy plus task-level automation data.

3 min read

MiniMax launches Speech 2.8 with sound tags and voice cloning

MiniMax introduced Speech 2.8, a synthetic speech model with native sound tags for breaths and hesitations, voice cloning from a 10-second sample, and processing intended to reduce noise and distortion. The company also says it improved cross-lingual speech for Mandarin and Japanese and made the model available through its platform and Audio product.

2 min read

DeepSeek Releases V4-Flash-Vision-Exp Multimodal Model

DeepSeek released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model on its API platform. The company says it matches DeepSeek-V4-Flash on text tasks while adding vision, and claims a major jump in multimodal agent benchmark performance over V4-Flash, near Opus-4.8. DeepSeek also launched a free Files API and DeepSeek Harness 0.1.1.

1 min read

What is a context window, and why does it run out?

A context window is the maximum amount of text, measured in tokens, that a model can consider in a single request. It holds the system instructions, the conversation so far, any documents you paste in, and the answer being generated. When the total exceeds the limit, something has to be dropped or summarised.

4 min read