Skip to content
inference.academy

Explainer

Sparse attention still has a storage bill

A query can read a small part of a long context without throwing the rest away. That reduces attention work, but it leaves the cache waiting in memory for the next query. Reading less and storing less move different parts of the serving bill.


The left grid is a session’s stored KV cache. The right grid is what one query reads. Change the query and the read budget first. Then lower the physically retained cache and watch which numbers finally move.

Original context

Read budget per query

Cache bytes per stored token

Stored for future queries

32,768 tokens / 512.0 MiB

Read by query 1

2,048 tokens / 32.0 MiB

The read positions change. The retained positions do not.

Cache per session
512.0 MiB
100% of the original 512.0 MiB
Cache read per decode step
32.0 MiB
6.3% of the retained cache
Sessions in 24 GiB of cache space
48
48 with the full cache; weights already excluded
Dense read of the retained cache512.0 MiB
Sparse read of the same cache32.0 MiB

Reading 2,048 tokens does not evict the other 30,720. Change the read budget: the read bar moves, while storage and session capacity stay fixed.

A ledger, not a model trace. Each cell groups 256 positions; both totals are exact, but routes and retained sets are invented. Bytes per token are illustrative totals across caching layers, not GLM measurements. Indexer state, weights, new writes, workspace and allocator overhead are excluded. Read bytes do not predict latency or accuracy.

A read budget is not a storage budget

With the default settings, 32,768 stored positions at 16 KiB per token occupy 512 MiB. Reading 2,048 positions fetches an idealised 32 MiB of cache state per decode step. It is one sixteenth of the read traffic, but still 512 MiB of storage. A 24 GiB cache budget holds 48 such sessions under either read policy.

Increase context to 131,072 tokens with the same read budget. The read total stays at 32 MiB while storage grows to 2 GiB. The cache budget now holds only 12 sessions. Fixed per-query attention does not imply fixed per-session memory.

Lower physical retention to 50%. Now storage falls, so session capacity rises. If more than 2,048 entries remain, the sparse query can still read 2,048. Compression has moved the storage ledger without changing this query’s read count.

Not selected does not mean safe to delete

A sparse indexer routes a particular query to a subset of positions. A different query can need a different subset. A position with no observed attention might have been excluded by routing rather than found irrelevant. Treating every unselected position as disposable confuses today’s route with tomorrow’s requirements.

Ramp’s GLM 5.3 experiments describe a maximum of 2,048 positions per query, with selected positions shared within groups of layers. Their reconstruction scoring uses dense attention so excluded positions can still receive a score. The lab illustrates that distinction, without pretending to reproduce the learned indexer.

Three separate ways to spend less

Read fewer positions. Sparse selection changes which cached positions contribute to a query. The positions can remain available for later queries.

Store fewer positions. Eviction or compaction removes entries or builds a smaller representation. Future queries must work with what survives.

Use fewer bytes per position. Quantization and latent representations can reduce the size of stored state. That changes storage even if the position count stays fixed.

Architectures can combine these. Sparse attention is a family of designs; some also store compressed summaries or bound their cache with a sliding window. This page isolates sparse selection over a stored token cache, rather than claiming that every sparse model has the same memory behaviour.

What the bytes do not tell you

The read count times bytes per token is an ideal cache payload. It excludes indexer state, selection overhead, weights, workspace, repeated reads and cache-line effects. A gather kernel does not necessarily deliver the same bandwidth as a contiguous read. Smaller payloads alone do not predict token latency.

The coloured routes and retained sets are invented, and no accuracy score is attached to them. Real selection has to preserve information across heads, layers and future queries. The point of changing the query here is to see that the route moves while storage remains.


Read next