Why Nvidia's Inference GPU Skips HBM for GDDR7

 "More inference means more HBM demand" is a common shorthand. Nvidia's own inference-only GPU says otherwise: Rubin CPX, unveiled September 9, 2025, carries no HBM at all. It runs on 128GB of GDDR7 — the same memory family used in gaming graphics cards.

KEY TAKEAWAYS

1. Rubin CPX, Nvidia's inference-specialized GPU announced Sep 9, 2025, uses 128GB of GDDR7 instead of HBM4. Shipments are planned for late 2026.

2. Its estimated memory bandwidth is roughly 2.1TB/s — about 1/10th of Rubin's 22TB/s HBM4 bandwidth. That's a design choice, not a shortfall: the workload it targets needs less memory bandwidth.

3. Nvidia splits inference into two stages — prefill and decode — and now builds separate chips for each. Prefill runs on Rubin CPX; decode runs on HBM4-equipped Rubin.

Estimated memory bandwidth by chip. Rubin CPX figure is back-calculated from rack-level totals, not a disclosed per-chip spec.

What Rubin CPX actually is

Nvidia describes Rubin CPX as a new class of GPU built for large-context processing — workloads like million-token codebases or long-form video, where the model has to hold and reason over huge amounts of input.

The disclosed specs: 128GB of GDDR7 memory (not HBM), up to 30 petaflops of NVFP4 compute, and built-in video encode/decode. Shipping is planned for late 2026. At the rack level, Nvidia's Vera Rubin NVL144 CPX is positioned at 8 exaflops, 100TB of fast memory, and 1.7PB/s of memory bandwidth — a 7.5x AI performance claim versus the GB300 NVL72. That comparison is company-provided, not an independent benchmark.

Inference is really two different jobs

Generating a response from a language model happens in two stages, and they behave nothing alike.

Prefill reads the entire input at once — feed it a million tokens, and it processes them in one large batch. The compute load is enormous. Memory traffic, relative to that compute load, is comparatively light.

Decode generates the response one token at a time. Every single token requires re-reading everything computed so far. Compute load per step is small. Memory traffic is constant and heavy.

Run both stages on the same GPU, and one of the two resources sits idle much of the time. Nvidia's answer is to split them across two different chips: Rubin CPX handles prefill, and HBM4-equipped Rubin handles decode.

Conceptual illustration based on Nvidia's stated design rationale for Rubin CPX (Sep 2025). Not a benchmarked measurement.

So it's GDDR7

GDDR7 is JEDEC's graphics DRAM standard (JESD239), finalized in March 2024. It runs at 32Gbps per pin and up to 192GB/s per chip — double GDDR6's bandwidth, achieved by switching signaling to PAM3.

Those numbers look modest next to HBM. But GDDR is a part you can mount with standard surface-mount assembly, right on the board next to the GPU. HBM requires through-silicon vias, a stacked die, and a silicon interposer to connect it to the GPU — a fundamentally more expensive and capacity-constrained process.

The Next Platform summed up the design choice in a September 2025 piece as being aimed at "lower cost and higher volume." Using expensive HBM where the workload doesn't need the bandwidth is, by that logic, waste.

JEDEC JESD239 (GDDR7) vs. JESD270-4 (HBM4) standard specifications.

What this means for memory demand

Two things pull in opposite directions here. First: "more inference equals more HBM demand" doesn't hold as a clean equation anymore. As the prefill share of inference workloads grows, that share of memory spend flows to GDDR, not HBM. HBM demand tracks the decode share more tightly.

Second, the counter-argument: if prefill gets cheaper, total inference volume can grow — and if decode volume grows in step, HBM demand could rise anyway. Which effect wins is an open question; there isn't data yet to settle it.

What is settled is this: graphics DRAM has entered the data center parts list. Until now, GDDR belonged almost entirely to gaming graphics cards.

Company-stated AI performance comparison, NVL144 CPX vs. GB300 NVL72. Not an independent benchmark.

What I actually watch

CheckpointWhy it matters
Whether Rubin CPX actually ships in late 2026The base Rubin GPU has already slipped before; CPX is still at the announcement stage.
GDDR7 supplier mixAll three memory makers produce graphics DRAM, but the share structure differs from HBM. Watch how data-center volume shifts that mix.
Whether prefill/decode splitting becomes an industry standardIf other accelerator makers copy the architecture, GDDR's data-center entry becomes structural, not a one-off.
Graphics DRAM revenue lines in earnings callsGDDR has been too small to break out separately until now. A standalone mention would be a signal.

Value chain read-through

SegmentWhat's soldIndicator to track
Inference front-endGDDR7Actual Rubin CPX shipments, adoption rate
Inference back-endHBM4Growth rate of decode-stage volume
SubstrateHigh-speed signaling PCBsGDDR7-compliant spec adoption
Back-end packagingStandard packagingEquipment burden vs. HBM packaging
SystemRack-scale integrated productsAdoption rate of CPX-equipped racks

Risks to this view

• The 2.1TB/s bandwidth figure for Rubin CPX is back-calculated from rack-level totals. The actual per-chip spec, once disclosed, could differ.

• The 7.5x performance claim is company-disclosed under specific test conditions. Real-world workload results may vary.

• Splitting prefill and decode means customers need to buy and operate two different types of hardware. Added operational complexity could slow adoption.

• If data-center GDDR7 volume grows large enough, it could tighten supply for consumer graphics cards and affect their pricing.

Next: how far apart Nvidia's announcement dates and actual shipment dates really are — we counted it across five GPU generations.

Sources: Nvidia Rubin CPX announcement (Sep 9, 2025); JEDEC GDDR7 standard JESD239 press release (Mar 5, 2024); The Next Platform (Sep 11, 2025); The Register (Sep 10, 2025); Nvidia Rubin disclosed specifications (Jan 2026).

Disclaimer: This post is for informational and educational purposes only. It does not constitute investment advice or a recommendation to buy or sell any security. All investment decisions are your own responsibility.

Comments

Popular posts from this blog

Korea's August Chip Exports Hit a Record $46.7B. Volume Moved Too

DDR4 Costs More Than DDR5 — Unless You're Actually Buying It