Explainer
Editing context can make the next turn more expensive
A shorter history should be cheaper to read. But a serving engine does not cache text; it caches the states that text produced in a particular history. Change the middle and the words after it can stay identical while their states become wrong. The next turn pays to rebuild them.
One fully cached history, one edit, one next turn. The green part can keep its states. Orange needs Prefill. Move the edit toward the start, then toward the end. The same replacement text can create very different amounts of work.
Cached history
Replacement text
- History after the edit
- 24,832
- 7,936 fewer tokens than before
- Exact prefix still reusable
- 6,144
- Tokens before the first change
- Unchanged text recomputed
- 18,432
- The edit changed its preceding context
The text is shorter, but this turn must prefill 18,688 tokens. Only 256 are new. The other 18,432 survived the edit and still need new states.
Same words, different states
In a causal transformer, later tokens can attend to earlier ones. Their hidden states, and therefore keys and values in deeper layers, depend on what came before. Position can matter too. A matching sequence after an edit is not enough to establish a matching KV cache.
The reusable part is the longest unchanged prefix. A replacement at the beginning leaves none of the old prompt prefix reusable. A replacement at the end leaves almost all of it reusable. Merely appending preserves the whole existing history. This is why moving a frequently updated scratchpad from the front to the end can change the amount of prefill without changing which information is retained.
Try “Edit near the start.” Replacing 8,192 old tokens with a 256-token note shrinks the 32,768-token history to 24,832. Exact reuse still needs to prefill all 24,832. Move the same edit to the end and it needs to prefill only 256. The text reduction is identical; the cache result is not.
The cost arrives now; the benefit can arrive later
Removing irrelevant context can improve an agent’s work and reduce what future turns carry. That benefit does not erase the immediate rebuild. A useful ledger records both: how much is recomputed at this edit, and what each later turn saves while the shorter history stays useful.
Tokens are the counting unit here, not a stopwatch. Prefilling one long suffix is not necessarily as slow as prefilling the same number of tokens in many small requests. Attention work, batching, kernels and hardware all change the time. The lab isolates which states can be reused before those questions enter.
Suffix reuse changes the contract
Enable the experimental comparison. Context Language Models proposes retaining states for surviving spans after an edit and adjusting their positional encoding. That avoids rebuilding the suffix, but its states still carry information from the old prefix. It approximates a fresh forward pass.
The paper reports matching task accuracy in its serving evaluation. That is evidence on the tested tasks, not proof of identical states or an exact output distribution. Its implementation also limits how many spans can be moved and can fall back when private cache slots cannot be allocated. The drawing assumes one suffix span and enough capacity.
What this model leaves out
The prompt starts fully resident, with an ideal token-level cache. Real serving engines reuse at block boundaries, evict histories, and format messages in ways that can change the effective prefix. Some strip earlier reasoning tokens. Those can create extra invalidation even when the visible conversation looks unchanged.
The lab also makes no decision about what an agent should forget. A cheap edit that removes a necessary fact is still a bad edit. Work avoided and task quality need separate measurements.
Read next
- explainers/
- sources/