• Tech Support ⤴
  • Projects
  • Services
    • AI Development
    • UI/UX Design
    • Web Development
    • Technology Support
    • Mobile App Development
    • Banking ATM Interfaces
    • Process Automation
    • Security Auditing
    • Local AI Servers
  • odoo ERP
get in touchStart with Eva
logo
Tech Support ⤴
Projects
Services
AI DevelopmentUI/UX DesignWeb DevelopmentTechnology SupportMobile App DevelopmentBanking ATM InterfacesProcess AutomationSecurity AuditingLocal AI Servers
odoo ERP
get in touchStart with Eva
Loading…
logo

Transforming businesses through AI-powered digital innovation and creative excellence.

Quick Links

BlogAinexProjectsContact us

Contact Us

pinDubai Digital Park, A5, DTEC - Silicon Oasisemail[email protected]phone+971 55 7538087
© 2026 aratech. All rights reserved.
Privacy PolicyTerms of ServiceCookie Policy
Home / Blog / The KV-Cache Endgame: How DeepSeek and Xiaomi Are Making Cheaper Models Smarter

The KV-Cache Endgame: How DeepSeek and Xiaomi Are Making Cheaper Models Smarter

DeepSeek-V4.1-Flash and Xiaomi's HySparse2 attack the same bottleneck from opposite directions: encode the KV cache once and quantize it hard, or read only what you will use. Agent-focused sparse attention is going mainstream, and cheap models just got competitive on real agent benchmarks.

- 6 min read

Key Takeaways

ExpandCollapse
  • - KV-cache memory bandwidth, not compute, is now the core inference bottleneck: a 1M-token context at full precision costs ~389 GB.
  • - DeepSeek-V4.1-Flash uses incremental KV encoding and FP4 (E2M1) caches to cut cache size 437x vs V1, reaching ~890 B per token and a ~0.9 GB 1M-token working set.
  • - Xiaomi HySparse2 shrinks what is read per token with elastic per-head sparsity, interleaved top-K + local retrieval, and shared grouped-query KV: 2.69 GB at 1M tokens, 5.02x less prefill compute vs hybrid sliding-window attention.
  • - Cache-optimized cheap models now rival frontier ones on agent benchmarks: V4.1-Flash (90.6) edges Claude Opus 5 (89.1) and GPT-5.6 (88.8) on Terminal-Bench 2.1.
  • - Million-token recall claims remain independently unverified beyond ~256K. Treat 1M recall as an architecture claim until third parties replicate it.
Line charts showing KV-cache footprints of DeepSeek and Xiaomi models and agent benchmark scores

The most expensive resource in AI right now is not compute. It is memory bandwidth, and nearly all of it is spent moving one thing: the KV cache. Two releases in the last few months attack that bottleneck from opposite engineering directions, DeepSeek's V4.1-Flash and Xiaomi's HySparse2, and together they sketch what the next era of inference looks like: million-token agents that fit in gigabytes, not terabytes.

1. First, size the problem

In a transformer's attention layer, generating one token requires reading the key/value vectors of every preceding token. At full attention with 61 layers (DeepSeek-V1), 128 attention heads and fine-grained head-dim of 128, a single token's cache at near-lossless precision costs roughly 380 KB across all layers. Multiply by context:

  • 32K tokens: ~12 GB of KV cache
  • 128K tokens: ~49 GB
  • 1M tokens: ~389 GB

That last number is why "1M context" was marketing until this year. It is not a model capability problem, it is a memory subsystem problem: even at 8 TB/s HBM bandwidth, streaming 389 GB per token caps you at roughly 20 tokens/second, before any thinking.

2. DeepSeek-V4.1-Flash: encode the cache once, quantize it hard

DeepSeek's lineage collapses that table:

ModelCache per token1M-token KV cache
DeepSeek-V1 (MLA, near-lossless)~389 KB~389 GB
DeepSeek-V4-Flash~3.5 KB (437x smaller)~3.5 GB
DeepSeek-V4.1-Flash~890 B (437x smaller)~0.89 GB

Two mechanisms do the work.

Incremental KV encoding. Earlier layers in the stack act as a memory hierarchy: encoding is applied only in a small final window of layers instead of recalculating the full stack per token. Context is built up by a tiny "diff" per token, with metadata structures indexing what has already been encoded, so repeated system prompts and reused tool outputs are encoded once, not per turn.

FP4 (E2M1) KV cache with FP8 metadata. DeepSeek-V4-Flash shipped FP8 KV caches; V4.1-Flash goes further and stores most cached vectors in NVFP4-style E2M1 blocks, halving V4-Flash's footprint again. Sensitive or high-frequency elements stay in higher precision, which is where most of the recall quality is preserved.

The compounding result is roughly 437x smaller than V1 and 4x smaller than V4-Flash, which is what makes a million-token context a ~0.9 GB working set instead of a small data center.

The point of all of it is agents. Self-declared benchmarks where the harness context (system prompt, skills, tool definitions, memory) dwarfs the user's actual query are precisely the workload that incremental encoding and cache-reuse redundancy elimination target: an agent session is mostly re-sent boilerplate, and V4.1-Flash stops paying for it repeatedly.

3. Xiaomi HySparse2: do not read what you will not use

Separately, Xiaomi's agent lab published HySparse2 (arXiv 2609.26368), a hybrid sparse attention architecture built for an 80B MoE (A3B active) agent model. Where DeepSeek shrinks the cache per token, HySparse2 shrinks what is read per token. Key mechanics:

  • Elastic head-level sparsity. Instead of one global sparsity config, each head picks its own balance of local sliding-window attention vs semantic top-K retrieval. Retention heads can stay "hot" while syntactic heads go mostly local.
  • Multiple profilers. A small profiling pass categorizes tokens into local vs global semantic roles, then splits the KV stream into parallel lanes: a fine-grained local-dense lane and a coarse global lane, each with its own policy.
  • Interleaved top-K + local prefill. During prompt processing, attention alternates between global top-K selection blocks and dense local window blocks so page-level hits (code repos, long documents) do not evict the local context that generation actually continues from.
  • KV-selected metadata. A separately stored compact index over the quantized cache drives retrieval. Xiaomi's ablation shows correctness on their long-context agentic set drops from 27.32 to roughly with vs without it; the metadata lane is what keeps sparse recall from silently degrading on needle-type queries.
  • MQA-style shared KV. The cache is organized grouped-query (MQA-style) (per-layer shared KV across query heads), making the agent's cache about 4-4.5x smaller than a standard MLA layout at the same width.

At 1M tokens on the 80B model, with FP8 KV: HySparse2 holds 2.69 GB vs 6.72 GB for HySparse and 12.09 GB for hybrid sliding-window attention, while doing 5.02x less prefill compute and 2.92x less decode compute than hybrid SWA. Per Xiaomi's retrieval evals (RULER-v2-style at 32K): 58.45 for HySparse2 vs 35.74 for HySparse vs 32.61 for sliding-window.

Important honesty note: Xiaomi's million-token claims are on their own eval harness against synthetic retrieval sets like RedPajama variants; independent third-party verification of true 1M-token recall is still thin. The architecture is credible; the endpoint numbers deserve the usual skepticism until independent teams replicate them at depth beyond 256K.

4. Do the wins show up on real agent benchmarks?

DeepSeek publishes head-to-heads for V4.1-Flash at full reasoning effort:

BenchmarkDeepSeek-V4.1-FlashClaude Opus 5GPT-5.6
Terminal-Bench 2.190.689.188.8
DeepSWE v1.174.274.0n/a
CyberGym88.1n/an/a
Terminal-Bench 4.031.251.8n/a

There are two different stories in that table. On Terminal-Bench 2.1, essentially the current standard agentic coding/ops harness, V4.1-Flash is measured at 90.6, edging out Claude Opus 5 (89.1) and GPT-5.6 (88.8). On the harder Terminal-Bench 4.0, Opus 5 remains well ahead (51.8 vs 31.2), so it is not a blanket claim of frontier dominance. But the direction is unambiguous: a cache-optimized, heavily quantized model is now competitive on the benchmarks that actually matter to people deploying agents, not just on static Q&A.

5. What this means if you build with models

  1. KV cache engineering is becoming the real frontier of inference optimization. The cost of serving an agent is dominated by the harness context (the system prompt + MCP tool definitions + memory), which in one trace I measured was 93 tokens of user question versus ~100K tokens of harness overhead per request. Architectures that refuse to re-pay that every turn change the unit economics by orders of magnitude.
  2. "Cheap model" no longer implies "dumb model". The correlation between a model's price bracket and its agentic benchmark position is breaking: V4.1-Flash is positioned far below Opus 5 on cost but sits roughly level with it on Terminal-Bench 2.1/DeepSWE/CyberGym.
  3. Sparse attention went from paper to production. HySparse2 shows the recipe (per-head sparsity mix + interleaved selection + quantized cache + retrieval metadata) that will be a default in the next generation of open-weights agentic models. Watch which labs adopt variants of it first.
  4. Long-context verification still has to be earned. Claimed 1M-token recall should be treated as an architecture claim, not an anchor; demand independent replications at 512K+ before you build product paths that depend on needle-in-a-1M-stack retrieval.

If the last three years were about scaling parameters, the next three look like they will be about scaling what fits in memory per dollar. DeepSeek and Xiaomi just published the strongest evidence yet that this is a solvable engineering problem, and that the teams who treat the KV cache as a first-class storage subsystem, not a side effect of attention, will set the price curve everyone else follows.

References

  1. Xiaomi AI (2026). "HySparse2: Hybrid Sparse Attention for Million-Token Agents." arXiv:2609.26368. URL: https://arxiv.org/abs/2609.26368
  2. Xiaomi origin engineering notes on HySparse2: https://originshq.com/blog/hysparse2-hybrid-sparse-attention-agents/ (secondary analysis of the paper's ablations and 1M-token KV figures)
  3. MindStudio (2026). "DeepSeek-V4.1-Flash KV Cache Compression." https://www.mindstudio.ai/blog/deepseek-v4-1-flash-kv-cache-compression
  4. NVIDIA (2024). "NVIDIA Blackwell Architecture Technical Brief: NVFP4 (E2M1) quantization." https://resources.nvidia.com/en-us-blackwell-architecture
  5. Original cover video: "DeepSeek's New Model vs Xiaomi's Sparse Attention" https://youtu.be/85QP5JDZfQM (thumbnail used with credit)

Table of Contents

  • ↗1. First, size the problem
  • ↗2. DeepSeek-V4.1-Flash: encode the cache once, quantize it hard
  • ↗3. Xiaomi HySparse2: do not read what you will not use
  • ↗4. Do the wins show up on real agent benchmarks?
  • ↗5. What this means if you build with models
  • ↗References

Related Posts

Stylized Citrix NetScaler gateway appliance glowing over a dark circuit-board grid with a root shell prompt exposed

Citrix NetScaler CVE-2026-88771/88772: Exploited in the Wild

Two critical Citrix NetScaler zero-day RCEs were actively exploited before Citrix's September 27 bulletin. Here is what CVE-2026-88771 and CVE-2026-88772 actually do, the fixed builds, and why forensics must come before patching.

Necolas HamwiNecolas Hamwi
September 30, 2026 - 9 min read
Dark cyberpunk artwork of a breached glowing SharePoint-style server rack with neon purple and cyan light, symbolizing the actively exploited CVE-2026-65660 vulnerability

SharePoint CVE-2026-65660: Attackers Are Actively Breaching Unpatched Servers

Microsoft confirmed that attackers are actively exploiting CVE-2026-65660, a SharePoint Server deserialization RCE chained with anonymous-access misconfigurations to drop web shells. CISA added it to KEV on September 25 and federal agencies must patch by September 28. Here is what the exploit chain looks like and a prioritized defense checklist.

Necolas HamwiNecolas Hamwi
September 29, 2026 - 7 min read
Dark cyberpunk illustration of a webmail server database under SQL injection attack, with neon purple and cyan circuit lines

Roundcube's Forgotten Plugin: A Four-Month-Old SQL Injection Is Now Running in the Wild

Roundcube Webmail's virtuser_query plugin carries CVE-2026-48842, a pre-authentication SQL injection that was patched back in May 2026. On September 24, Canada's Cyber Centre confirmed attackers are exploiting it in the wild, and any unpatched webmail server is an open door into the database behind it.

Necolas HamwiNecolas Hamwi
September 25, 2026 - 7 min read