DeepSeek V4.1 Cut KV Cache 4x. Parameters Nearly Doubled.

 DeepSeek V4.1 Flash shrinks its global KV cache to 890 bytes per token, about a quarter of V4 Flash. The same model card shows backbone parameters nearly doubling. For memory investors the net effect is a shift between memory tiers, not a clean cut in HBM demand.

KEY TAKEAWAYS

  1. The model card puts global KV cache at 890 bytes per token, roughly 1/4 of V4 Flash. Persistent KV stored on SSD falls to about 1/8.
  2. Backbone parameters rose from 284B to 552B (1.9x), and a 196B-parameter Engram memory module is listed separately.
  3. Illustrative: the same 2.3 TB of HBM holds 1,968 one-million-token sessions instead of 567. The saving shows up first as capacity for more sessions.

What the model card says, and what it does not

Chart 1. Global KV cache per token; V1 and V4 values derived from the card's ratios (log scale).

DeepSeek describes V4.1 Flash as a multimodal mixture-of-experts model with 552B backbone parameters and a context window of up to 1M tokens. Global KV cache is 890 bytes per token, about 4x smaller than V4 Flash and about 437x smaller than DeepSeek V1.

Two details got lost in the chatter. First, the card never says "HBM". Global KV is what normally sits in accelerator memory during decoding, so the reading is reasonable, but it is an interpretation. Second, the 1/4 and 1/8 figures refer to different things: global KV and persistent KV respectively. Neither is total memory use.

How the cache got smaller

KV cache stores attention keys and values for every token already processed. It grows linearly with context and is read again for each generated token, so it costs both capacity and bandwidth. At 1M tokens, V4.1 needs about 0.89 GB per session for global KV, against about 3.56 GB (derived) for V4 Flash.

Three design choices do the work. CSA2 shares KV and sparse-attention indices across layers. Main KV is cached in FP4. And a causal encoder-decoder layout projects the decoder's global KV once from the encoder's final hidden states. Sliding-window KV is rebuilt by replaying recent tokens instead of being persisted to SSD.

What went up

Chart 2. Backbone parameters nearly doubled while active parameters per token stayed in the same range.

Backbone parameters went from 284B to 552B. Active parameters per token are 8B in prefill and 16B in decode, close to V4 Flash's 13B, but keeping weights resident scales with the total, not the active count.

The card also lists a 196B-parameter Engram conditional memory accessed by token lookup. It does not say whether that sits inside the 552B or where it is served from. DeepSeek's January Engram paper, however, reported offloading a 100B-parameter table entirely to host DRAM with under 3% throughput loss. That is capacity that can live in ordinary server DRAM rather than HBM.

Where the saving goes



Chart 3. Illustrative calculation: 8 x 288 GB, 1 byte per weight, global KV only.

Take eight accelerators with 288 GB each, load weights at one byte per parameter and give the rest to global KV. V4 Flash fits 567 one-million-token sessions. V4.1 Flash fits 1,968. Illustrative calculation. Not actual company figures.

So the first-order effect is more long-context sessions per unit of HBM. Whether total HBM demand falls depends on how usage responds to cheaper sessions, and the model card has nothing to say about that. Efficiency gains in AI have often been followed by more spending, not less, which is why I would not draw a demand forecast from an architecture document.

Put by tier: HBM sees smaller KV but heavier weights. Host DRAM gains a possible home for Engram-style tables. SSD loses some persistent-KV volume. That is a reshuffle, and each tier has different suppliers and pricing.

Why DeepSeek optimized for input-heavy work

The card is explicit about the target: activating only 8B parameters during prefill improves cost efficiency for input-heavy agentic workloads. Coding agents and document agents read far more tokens than they write, often re-reading large repositories or files across many steps. That is exactly where KV size and prefill compute dominate the bill.

Seen that way, V4.1 is less a memory-saving model than a model built to make long, input-heavy sessions cheap enough to run at scale. If it works, the natural response is more such sessions, which is the channel through which memory demand could rise rather than fall. The card gives the mechanism but no usage data, so the direction stays an open question.

What I actually watch

CheckpointWhy it matters
Serving layout in the technical reportWhere weights, Engram and KV actually live
DeepSeek long-context and cache-hit pricingA read on usage elasticity
Whether other labs adopt similar KV compressionOne model vs an industry pattern
Context-length mix at inference providersWhether 1M-token sessions become common

Value chain read-through

SegmentLinkSignal
HBMLess KV per session, more weight capacityHBM per accelerator
Server DRAM, CXLOffloaded lookup tablesDRAM per CPU, CXL adoption
Enterprise SSDLess persistent KVInference SSD orders, context-caching prices

Risks to this view

  • This is one model from one lab and does not represent industry inference memory demand.
  • V4 Flash and V1 figures are back-calculated from approximate ratios in the card.
  • The session illustration simplifies weight precision, Engram placement and activation buffers.
  • I am deliberately not forecasting total HBM demand; the source does not support it.

KV compression moves memory between tiers more than it removes it. Next, I check whether Murata's MLCC part discontinuations are a supply cut or a product-mix shift, using Murata's own IR numbers.

Sources: DeepSeek V4.1 Flash model card on Hugging Face (Sep 2026); DeepSeek Engram paper, arXiv 2601.07372 (Jan 2026); author back-calculations and illustrative math.

Disclaimer: This post is for informational and educational purposes only. It does not constitute investment advice or a recommendation to buy or sell any security. All investment decisions are your own responsibility.

Comments

Popular posts from this blog

Why Nvidia's Inference GPU Skips HBM for GDDR7

Korea's August Chip Exports Hit a Record $46.7B. Volume Moved Too

DDR4 Costs More Than DDR5 — Unless You're Actually Buying It