HBM Is Not Enough: Where to Put 32TB of KV Cache

 The fastest-growing memory consumer in AI inference is not model weights. It is the KV cache, and it scales with context length and concurrency while HBM capacity stays bounded by what fits on the GPU package.

The industry answer is tiering. Marvell's recent product disclosures illustrate the architecture well — but they arrived in two separate announcements five months apart, at different stages of readiness. Sorting out which is which turns out to matter more than any single specification.

KEY TAKEAWAYS

1. Marvell states a CXL switch reaching up to 48TB of shared memory, and an optical fabric offloading up to 32TB of warm KV cache across up to 50 meters. All figures are company-stated.

2. The "up to 2-3x higher token throughput" claim is a target. Measurement conditions — model, context length, batch size — have not been disclosed.

3. Nothing here is in production. The CXL switch was slated to sample in Q3 2026, the SSD controller in Q4 2026, and no sampling date has been given for the optical fabric.

Why the KV cache creates a tiering problem

A KV cache stores the key and value vectors for tokens a transformer has already processed, so the model does not recompute them. It is what makes long-context inference tractable, and it grows roughly in proportion to context length multiplied by concurrent requests.

HBM is the fastest place to keep it and the most constrained. When the cache exceeds what HBM holds, the system either recomputes or waits. Neither is free, which is what creates room for an intermediate tier that is slower than HBM but far faster than pulling from storage.

Each tier trades latency against capacity, and each product targets a different row.

Three products, two announcements, three stages

This is the part that most coverage collapses. Marvell disclosed these separately, and they are not equally far along.

ProductRoleAnnouncedStage
Structera S 30260CXL switch, rack-scale poolingMarch 17, 2026, OFCSampling expected Q3 2026
Bravera SC6PCIe 6.0 SSD controllerAugust 4, 2026, FMSSampling expected Q4 2026
Photonic FabricMulti-rack optical shared memoryAugust 4, 2026, FMSNo sampling date disclosed
A live demo is not sampling, and sampling is not production.

The CXL switch was demonstrated live at OFC in March, with sampling to follow. The optical fabric was introduced in August as foundational elements of an architecture. Development complete, sampling, qualified and shipping are four different things, and the gap between the first and the last is where most forecasting errors live.

Expansion and sharing are different problems

CXL lets processors access memory coherently over the PCIe physical layer. It splits into two distinct use cases: adding DRAM capacity to a single server, and letting multiple hosts share a pool. Only the second needs a switch.

Per Marvell's stated specifications, the Structera S 30260 is a 260-lane device supporting CXL 3.0 with aggregate bandwidth up to 4TB/s. The company describes configurations connecting 16 or 32 CPUs or GPUs to as much as 48TB of shared memory at under 460ns round-trip latency, supporting both DDR5 and DDR4.

Optical solves distance, not latency

Where the CXL switch works inside a rack, the optical fabric targets the space between racks. Marvell describes a shared memory tier spanning multiple XPUs and racks at up to 50 meters, holding up to 32TB of warm KV cache offload.

The word doing the work is "warm." Active cache still has to live in HBM. What migrates is data likely to be needed again soon but not right now — the argument being that retrieving it from an optical shared tier beats reloading it from storage.

How far the performance numbers are verified

Marvell states the architecture can deliver up to 2-3x higher token throughput within existing data center footprints and power envelopes. Two qualifiers deserve attention. It is an "up to" figure, and the measurement conditions are not public.

The same applies to the switch benchmarks: up to 16x better scalability than local memory, 72% lower latency than RDMA-based pooling, up to 4.8x higher inference throughput in GPU configurations, and an 82.7% reduction in time to first token. These are vendor benchmarks. They are specific and plausible, and they are not independently verified — cite them as company figures or not at all.

Specific numbers, undisclosed test conditions, single source.

Value chain read-through

The most common misreading is that moving KV cache off HBM reduces HBM demand. The architecture does not support that conclusion. Adding tiers increases total memory content per rack, and the active cache still sits in HBM. What changes is the mix, not the sum.

SegmentIf tiering proceedsWhat I watch
HBMActive cache demand holds; capacity pressure partly shiftsHBM content per GPU over time
Server DRAM (DDR5)Feeds CXL modules and poolsCXL-capable module adoption announcements
NAND and SSD controllersTracks cold tier expansionPCIe 6.0 controller sampling and design wins
Optical componentsTracks inter-rack link demandWhether a sampling date is ever disclosed

Risks to this view

• Stage confusion. Expected sampling is not production. None of these three is shipping in volume.

• Undisclosed conditions. The 2-3x and 4.8x figures cannot be transferred to other configurations without knowing how they were measured.

• Ecosystem dependency. CXL requires processors, operating systems and modules to support it together. A switch alone does not make a pool.

• Adoption risk. Announced architectures take years to reach data center designs, and some never arrive.

Tiering is not an attempt to replace HBM. It is an attempt to cover the range HBM cannot reach economically. When reading announcements in this space, establishing what stage a product is at is more informative than any headline specification. Next I will work through what CXL adoption would mean for server DRAM demand on a per-module capacity basis.

Sources: Marvell press releases dated March 17, 2026 (CXL switch) and August 4, 2026 (AI memory infrastructure), plus company product disclosures. Everything here is from public sources.

Disclaimer: This post is for informational and educational purposes only. It does not constitute investment advice or a recommendation to buy or sell any security. All investment decisions are your own responsibility.

Comments

Popular posts from this blog

Why Nvidia's Inference GPU Skips HBM for GDDR7

Korea's August Chip Exports Hit a Record $46.7B. Volume Moved Too

DDR4 Costs More Than DDR5 — Unless You're Actually Buying It