HBM Is Not Enough: Where to Put 32TB of KV Cache
The fastest-growing memory consumer in AI inference is not model weights. It is the KV cache, and it scales with context length and concurrency while HBM capacity stays bounded by what fits on the GPU package.
The industry answer is tiering. Marvell's recent product disclosures illustrate the architecture well — but they arrived in two separate announcements five months apart, at different stages of readiness. Sorting out which is which turns out to matter more than any single specification.
KEY TAKEAWAYS
1. Marvell states a CXL switch reaching up to 48TB of shared memory, and an optical fabric offloading up to 32TB of warm KV cache across up to 50 meters. All figures are company-stated.
2. The "up to 2-3x higher token throughput" claim is a target. Measurement conditions — model, context length, batch size — have not been disclosed.
3. Nothing here is in production. The CXL switch was slated to sample in Q3 2026, the SSD controller in Q4 2026, and no sampling date has been given for the optical fabric.
Why the KV cache creates a tiering problem
A KV cache stores the key and value vectors for tokens a transformer has already processed, so the model does not recompute them. It is what makes long-context inference tractable, and it grows roughly in proportion to context length multiplied by concurrent requests.
HBM is the fastest place to keep it and the most constrained. When the cache exceeds what HBM holds, the system either recomputes or waits. Neither is free, which is what creates room for an intermediate tier that is slower than HBM but far faster than pulling from storage.
Three products, two announcements, three stages
This is the part that most coverage collapses. Marvell disclosed these separately, and they are not equally far along.
| Product | Role | Announced | Stage |
|---|---|---|---|
| Structera S 30260 | CXL switch, rack-scale pooling | March 17, 2026, OFC | Sampling expected Q3 2026 |
| Bravera SC6 | PCIe 6.0 SSD controller | August 4, 2026, FMS | Sampling expected Q4 2026 |
| Photonic Fabric | Multi-rack optical shared memory | August 4, 2026, FMS | No sampling date disclosed |
The CXL switch was demonstrated live at OFC in March, with sampling to follow. The optical fabric was introduced in August as foundational elements of an architecture. Development complete, sampling, qualified and shipping are four different things, and the gap between the first and the last is where most forecasting errors live.
Expansion and sharing are different problems
CXL lets processors access memory coherently over the PCIe physical layer. It splits into two distinct use cases: adding DRAM capacity to a single server, and letting multiple hosts share a pool. Only the second needs a switch.
Per Marvell's stated specifications, the Structera S 30260 is a 260-lane device supporting CXL 3.0 with aggregate bandwidth up to 4TB/s. The company describes configurations connecting 16 or 32 CPUs or GPUs to as much as 48TB of shared memory at under 460ns round-trip latency, supporting both DDR5 and DDR4.
Optical solves distance, not latency
Where the CXL switch works inside a rack, the optical fabric targets the space between racks. Marvell describes a shared memory tier spanning multiple XPUs and racks at up to 50 meters, holding up to 32TB of warm KV cache offload.
The word doing the work is "warm." Active cache still has to live in HBM. What migrates is data likely to be needed again soon but not right now — the argument being that retrieving it from an optical shared tier beats reloading it from storage.
How far the performance numbers are verified
Marvell states the architecture can deliver up to 2-3x higher token throughput within existing data center footprints and power envelopes. Two qualifiers deserve attention. It is an "up to" figure, and the measurement conditions are not public.
The same applies to the switch benchmarks: up to 16x better scalability than local memory, 72% lower latency than RDMA-based pooling, up to 4.8x higher inference throughput in GPU configurations, and an 82.7% reduction in time to first token. These are vendor benchmarks. They are specific and plausible, and they are not independently verified — cite them as company figures or not at all.
Value chain read-through
The most common misreading is that moving KV cache off HBM reduces HBM demand. The architecture does not support that conclusion. Adding tiers increases total memory content per rack, and the active cache still sits in HBM. What changes is the mix, not the sum.
| Segment | If tiering proceeds | What I watch |
|---|---|---|
| HBM | Active cache demand holds; capacity pressure partly shifts | HBM content per GPU over time |
| Server DRAM (DDR5) | Feeds CXL modules and pools | CXL-capable module adoption announcements |
| NAND and SSD controllers | Tracks cold tier expansion | PCIe 6.0 controller sampling and design wins |
| Optical components | Tracks inter-rack link demand | Whether a sampling date is ever disclosed |
Risks to this view
• Stage confusion. Expected sampling is not production. None of these three is shipping in volume.
• Undisclosed conditions. The 2-3x and 4.8x figures cannot be transferred to other configurations without knowing how they were measured.
• Ecosystem dependency. CXL requires processors, operating systems and modules to support it together. A switch alone does not make a pool.
• Adoption risk. Announced architectures take years to reach data center designs, and some never arrive.
Tiering is not an attempt to replace HBM. It is an attempt to cover the range HBM cannot reach economically. When reading announcements in this space, establishing what stage a product is at is more informative than any headline specification. Next I will work through what CXL adoption would mean for server DRAM demand on a per-module capacity basis.
Sources: Marvell press releases dated March 17, 2026 (CXL switch) and August 4, 2026 (AI memory infrastructure), plus company product disclosures. Everything here is from public sources.
Comments
Post a Comment