KV Cache Compression and Memory Control (September 2026)¶
Version context: cache layouts and supported strategies depend on the model architecture, framework, backend, precision, and serving engine. Use the documentation and telemetry for the exact deployed combination. This page avoids universal compression ratios and performance claims.
A key-value cache stores attention keys and values already computed for prior tokens so autoregressive decoding does not recompute them at every generation step. It improves decoding efficiency, but retained cache memory can become the limiting resource for long contexts or concurrent requests.
What Determines Cache Size¶
For a decoder-only model with conventional attention, a planning estimate is:
The factor two represents keys and values. Grouped-query or multi-query attention can use fewer key/value heads than query heads, so use the model's actual number of KV heads. Encoder-decoder, sliding-window, linear-attention, and model-specific cache implementations may use a different layout.
def estimate_kv_cache_bytes(
*,
num_layers: int,
batch_size: int,
retained_tokens: int,
num_kv_heads: int,
head_dim: int,
bytes_per_element: int,
) -> int:
return (
2
* num_layers
* batch_size
* retained_tokens
* num_kv_heads
* head_dim
* bytes_per_element
)
Treat this as capacity planning, not an allocation receipt. Framework overhead, attention kernels, fragmentation, prefix reuse, graph capture, and parallelism also consume memory.
Diagnose Before Compressing¶
Collect a workload profile before choosing a strategy:
| Signal | Why it matters |
|---|---|
| Input and generated-token distributions | Determines retained sequence length |
| Concurrent active sequences | Multiplies cache allocation |
| Prefill versus decode latency | Separates prompt and token-generation bottlenecks |
| Cache hit/reuse ratio | Shows whether common prefixes are valuable |
| GPU memory, OOMs, fragmentation | Decides whether memory is the real constraint |
| Quality at target context length | Detects long-context regressions |
| Tenant/data scope | Prevents unsafe cache reuse |
Do not tune a cache on a short synthetic prompt and deploy it to a long-document or multi-user workload.
Choose a Cache Strategy by Constraint¶
| Strategy | Best when | Trade-off |
|---|---|---|
| Dynamic cache | Sequence lengths vary and simplicity matters | Allocation growth and less predictable shapes |
| Static/preallocated cache | Maximum length is known and framework support exists | Reserved memory can be wasted or cause OOM |
| Sliding-window cache | The model and task tolerate bounded recent context | Older context is no longer available to attention |
| Offloaded cache | Device memory is scarce and latency budget permits transfer | Host/device transfer can dominate latency |
| Quantized cache | Memory is the bottleneck and model/backend support it | Task-specific accuracy and kernel behavior must be measured |
| Paged/block-managed serving cache | Many concurrent sequences have uneven lengths | Engine-specific operational tuning |
| Prefix cache | Many requests share an identical immutable prefix | Requires exact cache keys and isolation controls |
Hugging Face documents dynamic, static, offloaded, and quantized cache variants. A serving engine may use a block or page allocator rather than a single contiguous tensor. Select the implementation supported by the deployed model and backend; do not copy flags from an unrelated engine.
Static Cache and Compile-Friendly Shapes¶
A static cache preallocates storage to a maximum cache length. In compatible Transformers configurations it can enable compile-friendly execution, but it raises the memory floor.
Set the maximum from observed request limits plus a justified margin. A large arbitrary maximum turns an optimization into an out-of-memory risk. Run an isolated soak test with the intended concurrency before enabling it for serving traffic.
Prefix Reuse Is a Data Boundary¶
Prefix reuse can avoid recomputing a common system prompt, document prefix, or shared context. It is safe only when the cache key includes every behaviorally relevant input:
{
"model_revision": "immutable-model-revision",
"tokenizer_revision": "tokenizer:sha256:...",
"chat_template_revision": "template:sha256:...",
"system_prompt_revision": "prompt:sha256:...",
"tenant_scope": "tenant_42",
"tool_schema_revision": "tools:sha256:..."
}
Never reuse cached state across tenants, users, permission scopes, or prompt revisions unless the policy explicitly permits that exact boundary. Cache invalidation here is a correctness and privacy issue, not just a performance issue.
Quantization and Compression¶
Quantized caches reduce bytes per stored value. The appropriate precision depends on model, attention architecture, implementation, context length, and task. A lower-memory cache may improve concurrency while degrading retrieval, structured output, or reasoning at longer contexts.
Validate candidates with:
- a fixed prompt and generated-token distribution;
- representative long-context and retrieval cases;
- structured-output and tool-use checks;
- quality comparison against the approved baseline;
- peak memory, p50/p95 latency, throughput, and error rate;
- a rollback condition if quality or OOM rate crosses a threshold.
Compression research can suggest candidates, but paper numbers are not a production capacity plan.
Serving Operations¶
Separate two concerns:
per-request correctness: prompt, cache key, context limit, output validation
system capacity: batching, allocator pressure, memory headroom, admission control
For a multi-user service, admission control must consider cache capacity before accepting a request. Rejecting or queueing a request with an explicit capacity status is safer than accepting it and failing after partial work. Export cache memory, active sequences, hit/miss, evictions, OOMs, and latency as observability signals.
Verification Receipt¶
{
"candidate": "cache-strategy-revision",
"model_revision": "immutable-model-revision",
"workload_revision": "eval:sha256:...",
"context_distribution": "production-like-v3",
"quality_gate": "no material regression",
"capacity_gate": "no OOM at approved concurrency",
"rollback": "dynamic-cache-baseline"
}
A receipt makes clear whether a result applies to the intended model, engine, workload, and hardware. Re-run it after changing model weights, tokenizer, context policy, kernel/backend, or serving engine.
Gotchas¶
- KV-head count is not always attention-head count. GQA and MQA change the memory calculation. Fix: read the actual model configuration before estimating capacity.
- Static allocation can fail before the first token. A large maximum reserves memory even for small requests. Fix: size it from a measured limit and test concurrency.
- A cache key can accidentally cross a data boundary. Similar-looking prompts may have different tenant, tools, or permissions. Fix: include revisions and scope in the key and fail closed on a mismatch.
- Quantization gains are task-dependent. A cache that preserves short-form perplexity can fail long-context retrieval. Fix: validate target tasks at the deployed context length.
- Caching is an inference mechanism. Enabling it in training can cause unexpected behavior. Fix: follow the framework's training guidance and keep training and serving configurations separate.
Sources¶
- Hugging Face cache strategies
- Hugging Face cache explanation
- Hugging Face inference optimization
- vLLM documentation
- Transformers generation utilities