The AiEdge Newsletter

The AiEdge Newsletter

How DeepSeek V4.1 Flash Separates Reading and Generation

How a shorter input path, shared attention memory, and selective reconstruction change the cost of long conversations.

Damien Benveniste's avatar
Damien Benveniste
Oct 07, 2026
∙ Paid

DeepSeek V4.1 Flash generates text one token at a time. A token is a piece of text, often a word or part of one. To choose what comes next, the model transforms the text already available through successive processing stages called layers. Its unusual feature is that most supplied text can take a shorter path through those stages than a token being used to generate the next one.

DeepSeek's technical report motivates this design with agents that repeatedly receive long tool results and accumulated history, making the cost of reading and retaining that information a major part of running the model. Why should preparing old text for later use require the same computation as deciding what to write next?

Most supplied positions pass through the first 20 layers to prepare memory. The final 128 also prepare recent context in the next 20 layers. Each chosen answer token uses both halves and earlier memory to predict the following token.
The shorter route prepares memory for later use. Recent input still enters the upper half before generation, and each chosen answer token uses both halves.

The design divides 40 layers into two halves. The first 20 prepare reusable memory of the input. Most input positions can stop there; a final slice of up to 128 positions also passes through the upper half to prepare its recent context. Once generation begins, each newly chosen token goes through both halves to predict the following token. The architecture is called a causal encoder-decoder, or CED. We will unpack those roles, then trace how sharing memory, selecting what to read, storing fewer bits, and rebuilding recent state change different costs.

The useful distinction is already visible: preparing a record of earlier text and using that record to choose a new token need not follow identical paths. The shorter input route matters most when there is much more new text to read than text to generate.

Table of contents

  • Why one causal stack reads and writes

  • Why agent workloads change the budget

  • Prefill, decode, and the cache between them

  • How CED changes the old-token dependency

  • Why generated tokens still use the full model

  • Why this is a different encoder-decoder

  • How CSA2 shares global history

  • Why sparse search needs a hierarchy

  • How the cache bytes add up

  • What bounded replay rebuilds

  • Putting the architecture together

  • What the quality evidence establishes

  • What this means for agent architectures

User's avatar

Continue reading this post for free, courtesy of Damien Benveniste.

Or purchase a paid subscription.
© 2026 AiEdge · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture