Skip to content
inference.academy

Explainer

When does compression pay for itself?

A smaller cache can make later work cheaper. Making that smaller cache can be expensive. Whether the trade pays depends on how much of each decode step can actually shrink, and how many times the compressed state is used before it is discarded.


One source cache is prepared once, then reused for independent continuations of equal length. Grey is the accumulated full-cache decode time. Green starts with a preparation charge and accumulates compressed-cache decode time. The intersection is where that charge is repaid.

All timings below are illustrative inputs. Replace them with measurements from the same workload and hardware.

Output tokens per use

Per-token overhead and the calculation

Compressed step = full step × (1 − cache share + cache share × retained fraction) + extra work. Every part outside the source-cache share stays unchanged. The compressed total adds preparation once.

Preparation moves the compressed line up; cheaper decode steps make its slope shallowerAt 32 uses, full cache costs 163.8 s and compressed cache including preparation costs 142.2 s. First cheaper whole use: 19.0 s177 s354 s016324864Uses of the same prepared cache
Full cacheCompressed, including preparationYour reuse count
Estimated compressed step
13.70 ms
Full step: 20.00 ms
First cheaper whole use
19
256 output tokens each use
Time saved at this reuse count
21.6 s
32 uses, preparation included

Preparation has been repaid. Across 32 uses, full cache takes 163.8 sand compression takes 142.2 s, saving 21.6 s under these assumptions.

Linear sensitivity model for reusing one fixed source cache, with equal output lengths and one preparation charge. New output-cache growth, queueing, prefill, concurrency gains and quality changes are excluded. Cache-dependent time is assumed proportional to retained bytes; sparse reads and kernel overhead can break that assumption. These are elapsed-time estimates, not GPU-hours or API prices.

An 80% smaller cache is not an 80% faster step

A decode step also reads weights, computes projections, runs other layers and launches kernels. Shrinking the source cache leaves that work in place. In the lab, only the share of step time you assign to this cache scales with retained bytes.

With the illustrative defaults, a full step takes 20 ms and 40% of it scales with the source cache. Keeping 20% of that cache leaves 12 ms of unaffected work and 1.6 ms of cache-dependent work. Add 0.1 ms of extra work and the compressed step takes 13.7 ms. That saves 6.3 ms per output token, not 16 ms.

At 256 output tokens per use, the saving is 1.6128 seconds each use. A 30-second preparation charge first becomes cheaper on use 19. At 32 uses, the full path takes 163.84 seconds and the compressed path takes 142.2304, including preparation. These numbers are arithmetic from the inputs, not a measured speedup.

Preparation has several line items

Depending on the method, preparing a cache can include capturing scoring queries, selecting retained entries, fitting weights and values, rerunning earlier compressed layers, and arranging the result for serving. Measure the whole path. A timer on only the fast selector leaves the other work unpaid.

Ramp reports the selection times below for Attention Matching on a five-article, 90-question QuALITY slice across four runs. They are average key-selection times per article, not complete compression times or decode benchmarks.

Ramp’s reported selection costs and downstream accuracy
SelectionCache removedSelection timeAccuracy
HAK95%3.51 s73.7% ± 1.1pp
OMP95%21.49 s77.8% ± 1.9pp
OMP80%256.95 s83.9% ± 1.2pp

The full-context baseline is 91.8% ± 0.6pp. The 80% setting keeps more accuracy, but its selector is much slower. The preset loads 256.95 seconds as a lower bound on preparation. It keeps the decode inputs illustrative because the article does not report an end-to-end serving speedup.

Reuse pays a one-time cost, not a recurring loss

Increase the number of uses and a genuinely cheaper step eventually repays preparation. Set the cache-dependent share to zero, or make extra per-token work larger than the saving, and no amount of reuse pays it back. The compressed line has to have a shallower slope.

Reuse here means continuing from the same prepared source state. If each request needs a freshly compressed article, preparation is charged for each one and cannot be amortised across unrelated requests. A new setup charge is also needed whenever the prepared state is rebuilt. Identical text is not enough if the serving system cannot reuse the prepared cache.

Time is only one part of the decision

A smaller cache can admit more sessions or avoid eviction even when it does not reduce an individual decode step. Those capacity benefits are outside this curve. So are quality losses, which can make a cheaper answer less useful or trigger retries.

This model also freezes the source cache. Output tokens can build new state that grows during a continuation. Sparse attention may already cap the entries read, and real kernels do not necessarily scale linearly with stored bytes. To use the curve for a deployment decision, measure full and compressed work on the same tasks, hardware and batch conditions, check quality separately, and include every preparation stage.

The vertical axis is elapsed time. It becomes neither API cost nor GPU-hours automatically: preparation can use different hardware, overlap other work, or consume several devices. Keep that ledger separate.


Read next