DeepSeek V4.1 Flash generates text one token at a time. A token is a piece of text, often a word or part of one. To choose what comes next, the model transforms the text already available through successive processing stages called layers. Its unusual feature is that most supplied text can take a shorter path through those stages than a token being used to generate the next one.
DeepSeek's technical report motivates this design with agents that repeatedly receive long tool results and accumulated history, making the cost of reading and retaining that information a major part of running the model. Why should preparing old text for later use require the same computation as deciding what to write next?

The design divides 40 layers into two halves. The first 20 prepare reusable memory of the input. Most input positions can stop there; a final slice of up to 128 positions also passes through the upper half to prepare its recent context. Once generation begins, each newly chosen token goes through both halves to predict the following token. The architecture is called a causal encoder-decoder, or CED. We will unpack those roles, then trace how sharing memory, selecting what to read, storing fewer bits, and rebuilding recent state change different costs.
The useful distinction is already visible: preparing a record of earlier text and using that record to choose a new token need not follow identical paths. The shorter input route matters most when there is much more new text to read than text to generate.
Table of contents
Why one causal stack reads and writes
Why agent workloads change the budget
Prefill, decode, and the cache between them
How CED changes the old-token dependency
Why generated tokens still use the full model
Why this is a different encoder-decoder
How CSA2 shares global history
Why sparse search needs a hierarchy
How the cache bytes add up
What bounded replay rebuilds
Putting the architecture together
What the quality evidence establishes
What this means for agent architectures



