NVIDIA Unveils NVHBM: +30% Bandwidth, -15% Power vs HBM4e
NVIDIA has unveiled NVHBM, a custom HBM base-die technology for its NVLink Fusion partners. The headline claims: up to 30% more bandwidth and 15% lower power than standard HBM4e. Here is what it means for the memory value chain.
What was announced
On August 26, NVIDIA expanded its NVLink Fusion program with NVHBM, a custom high-bandwidth memory base-die technology for hyperscalers building semi-custom AI accelerators (XPUs). It is a component for custom chips, not a replacement for commodity HBM.
NVIDIA's claims versus standard HBM4e: up to 30% more memory bandwidth per stack, up to 15% lower HBM power consumption, and up to 25% more XPU compute-die area. Combined, the company says these translate into up to a 30% overall end-to-end performance increase per XPU.
What is confirmed and what is not:
- Confirmed (company announcement): the architecture, the figures above, and Amazon's Annapurna Labs as the first collaborator.
- Not disclosed: which memory makers will supply NVHBM (only "leading memory partners"), the production timeline, and pricing.
- Unverified reports: Taiwanese media reports from August 2025 pointing to a 3nm base die with small-batch trial production in 2H 2027 were not confirmed in this announcement.
1. The controller leaves the XPU
In conventional HBM designs, the memory controller and a wide PHY sit on the XPU compute die — meaning a meaningful share of the most expensive leading-edge silicon is spent on memory plumbing rather than compute.
NVHBM moves the controller into the base die of the 3D HBM stack and leaves only a narrow custom PHY on the XPU side. Per NVIDIA's technical blog, PHY and support area shrink by up to 67% versus the JEDEC HBM4e standard, and the simpler interposer routing yields up to 80% more usable silicon across the layout.
One inconsistency worth flagging: NVIDIA's summary table cites "up to 25% more compute die area," while the body of its technical blog says "up to a 30% increase in available main-die silicon." The baselines appear to differ (compute die vs. main die overall); the conservative reading is 25%.
2. Built for the inference bottleneck
Large-model inference is dominated by repeated reads of model weights and KV cache. However fast the compute engines are, they stall if HBM cannot feed them. A 30% per-stack bandwidth uplift directly raises utilization and token throughput in memory-bound phases.
The 15% power saving matters at scale. By NVIDIA's own illustrative calculation, in a hypothetical 1-gigawatt data center filled with 2,000W XPUs, the memory power saved is enough to run up to 15,000 additional XPUs — roughly 3% of the ~500,000 XPUs such a facility could host. With power now the binding constraint on AI infrastructure, that is not a small number.
3. Value-chain implications: who defines the base die?
Since the HBM4 generation, the base die has been migrating from memory makers' in-house designs toward logic foundries. NVHBM goes a step further: the design authority over the base die moves to the chip designer (NVIDIA).
For the three memory makers, the implications cut both ways:
- Positive: a single NVIDIA-validated implementation supplied by multiple memory vendors reduces the burden of designing and qualifying a custom base die for each customer, and pools custom-HBM demand into the NVLink Fusion ecosystem.
- Negative: if the base die standardizes on NVIDIA's design, memory makers lose room to differentiate through controller and logic features, pushing competition back toward DRAM core stacking and cost.
The first collaborator is notable, too: Amazon's Annapurna Labs. Its next-generation Trainium4 will support NVLink Fusion, with NVHBM adoption possible in subsequent designs. Hyperscalers that built custom silicon to reduce NVIDIA dependence are, at the memory subsystem level, adopting NVIDIA technology.
Checkpoints
- Memory partner disclosure: which of SK hynix, Samsung, and Micron sign on as validated NVHBM suppliers.
- Base-die foundry: process node and manufacturer (the reported 3nm / 2H 2027 trial production remains unconfirmed).
- Standard vs. custom mix: how commodity HBM4e volumes evolve alongside NVHBM custom volumes.
- NVLink Fusion adoption: which XPU designers follow Annapurna Labs.
Risks
- All figures are NVIDIA's "up to" claims; real-world workload results may differ.
- Memory suppliers and production timing are undisclosed, leaving commercialization uncertain.
- Hyperscalers committed to their own memory subsystems may limit adoption.
- Base-die standardization may conflict with memory makers' custom-HBM value-add strategies.
Relocating a memory controller looks like a small architectural shuffle, but on the question of who defines the HBM specification, it is a big move. We will revisit this when the memory partner list is disclosed.
Sources: NVIDIA blog and technical blog (Aug 26, 2026); Tom's Hardware; TrendForce (Aug 2025 reports).
Disclaimer: This post is for personal study and informational purposes only and is not a recommendation to buy or sell any security. Investment decisions are your own responsibility.





Comments
Post a Comment