The most expensive resource in AI right now is not compute. It is memory bandwidth, and nearly all of it is spent moving one thing: the KV cache. Two releases in the last few months attack that bottleneck from opposite engineering directions, DeepSeek's V4.1-Flash and Xiaomi's HySparse2, and together they sketch what the next era of inference looks like: million-token agents that fit in gigabytes, not terabytes.
1. First, size the problem
In a transformer's attention layer, generating one token requires reading the key/value vectors of every preceding token. At full attention with 61 layers (DeepSeek-V1), 128 attention heads and fine-grained head-dim of 128, a single token's cache at near-lossless precision costs roughly 380 KB across all layers. Multiply by context:
- 32K tokens: ~12 GB of KV cache
- 128K tokens: ~49 GB
- 1M tokens: ~389 GB
That last number is why "1M context" was marketing until this year. It is not a model capability problem, it is a memory subsystem problem: even at 8 TB/s HBM bandwidth, streaming 389 GB per token caps you at roughly 20 tokens/second, before any thinking.
2. DeepSeek-V4.1-Flash: encode the cache once, quantize it hard
DeepSeek's lineage collapses that table:
Two mechanisms do the work.
Incremental KV encoding. Earlier layers in the stack act as a memory hierarchy: encoding is applied only in a small final window of layers instead of recalculating the full stack per token. Context is built up by a tiny "diff" per token, with metadata structures indexing what has already been encoded, so repeated system prompts and reused tool outputs are encoded once, not per turn.
FP4 (E2M1) KV cache with FP8 metadata. DeepSeek-V4-Flash shipped FP8 KV caches; V4.1-Flash goes further and stores most cached vectors in NVFP4-style E2M1 blocks, halving V4-Flash's footprint again. Sensitive or high-frequency elements stay in higher precision, which is where most of the recall quality is preserved.
The compounding result is roughly 437x smaller than V1 and 4x smaller than V4-Flash, which is what makes a million-token context a ~0.9 GB working set instead of a small data center.
The point of all of it is agents. Self-declared benchmarks where the harness context (system prompt, skills, tool definitions, memory) dwarfs the user's actual query are precisely the workload that incremental encoding and cache-reuse redundancy elimination target: an agent session is mostly re-sent boilerplate, and V4.1-Flash stops paying for it repeatedly.
3. Xiaomi HySparse2: do not read what you will not use
Separately, Xiaomi's agent lab published HySparse2 (arXiv 2609.26368), a hybrid sparse attention architecture built for an 80B MoE (A3B active) agent model. Where DeepSeek shrinks the cache per token, HySparse2 shrinks what is read per token. Key mechanics:
- Elastic head-level sparsity. Instead of one global sparsity config, each head picks its own balance of local sliding-window attention vs semantic top-K retrieval. Retention heads can stay "hot" while syntactic heads go mostly local.
- Multiple profilers. A small profiling pass categorizes tokens into local vs global semantic roles, then splits the KV stream into parallel lanes: a fine-grained local-dense lane and a coarse global lane, each with its own policy.
- Interleaved top-K + local prefill. During prompt processing, attention alternates between global top-K selection blocks and dense local window blocks so page-level hits (code repos, long documents) do not evict the local context that generation actually continues from.
- KV-selected metadata. A separately stored compact index over the quantized cache drives retrieval. Xiaomi's ablation shows correctness on their long-context agentic set drops from 27.32 to roughly with vs without it; the metadata lane is what keeps sparse recall from silently degrading on needle-type queries.
- MQA-style shared KV. The cache is organized grouped-query (MQA-style) (per-layer shared KV across query heads), making the agent's cache about 4-4.5x smaller than a standard MLA layout at the same width.
At 1M tokens on the 80B model, with FP8 KV: HySparse2 holds 2.69 GB vs 6.72 GB for HySparse and 12.09 GB for hybrid sliding-window attention, while doing 5.02x less prefill compute and 2.92x less decode compute than hybrid SWA. Per Xiaomi's retrieval evals (RULER-v2-style at 32K): 58.45 for HySparse2 vs 35.74 for HySparse vs 32.61 for sliding-window.
Important honesty note: Xiaomi's million-token claims are on their own eval harness against synthetic retrieval sets like RedPajama variants; independent third-party verification of true 1M-token recall is still thin. The architecture is credible; the endpoint numbers deserve the usual skepticism until independent teams replicate them at depth beyond 256K.
4. Do the wins show up on real agent benchmarks?
DeepSeek publishes head-to-heads for V4.1-Flash at full reasoning effort:
There are two different stories in that table. On Terminal-Bench 2.1, essentially the current standard agentic coding/ops harness, V4.1-Flash is measured at 90.6, edging out Claude Opus 5 (89.1) and GPT-5.6 (88.8). On the harder Terminal-Bench 4.0, Opus 5 remains well ahead (51.8 vs 31.2), so it is not a blanket claim of frontier dominance. But the direction is unambiguous: a cache-optimized, heavily quantized model is now competitive on the benchmarks that actually matter to people deploying agents, not just on static Q&A.
5. What this means if you build with models
- KV cache engineering is becoming the real frontier of inference optimization. The cost of serving an agent is dominated by the harness context (the system prompt + MCP tool definitions + memory), which in one trace I measured was 93 tokens of user question versus ~100K tokens of harness overhead per request. Architectures that refuse to re-pay that every turn change the unit economics by orders of magnitude.
- "Cheap model" no longer implies "dumb model". The correlation between a model's price bracket and its agentic benchmark position is breaking: V4.1-Flash is positioned far below Opus 5 on cost but sits roughly level with it on Terminal-Bench 2.1/DeepSWE/CyberGym.
- Sparse attention went from paper to production. HySparse2 shows the recipe (per-head sparsity mix + interleaved selection + quantized cache + retrieval metadata) that will be a default in the next generation of open-weights agentic models. Watch which labs adopt variants of it first.
- Long-context verification still has to be earned. Claimed 1M-token recall should be treated as an architecture claim, not an anchor; demand independent replications at 512K+ before you build product paths that depend on needle-in-a-1M-stack retrieval.
If the last three years were about scaling parameters, the next three look like they will be about scaling what fits in memory per dollar. DeepSeek and Xiaomi just published the strongest evidence yet that this is a solvable engineering problem, and that the teams who treat the KV cache as a first-class storage subsystem, not a side effect of attention, will set the price curve everyone else follows.
References
- Xiaomi AI (2026). "HySparse2: Hybrid Sparse Attention for Million-Token Agents." arXiv:2609.26368. URL: https://arxiv.org/abs/2609.26368
- Xiaomi origin engineering notes on HySparse2: https://originshq.com/blog/hysparse2-hybrid-sparse-attention-agents/ (secondary analysis of the paper's ablations and 1M-token KV figures)
- MindStudio (2026). "DeepSeek-V4.1-Flash KV Cache Compression." https://www.mindstudio.ai/blog/deepseek-v4-1-flash-kv-cache-compression
- NVIDIA (2024). "NVIDIA Blackwell Architecture Technical Brief: NVFP4 (E2M1) quantization." https://resources.nvidia.com/en-us-blackwell-architecture
- Original cover video: "DeepSeek's New Model vs Xiaomi's Sparse Attention" https://youtu.be/85QP5JDZfQM (thumbnail used with credit)