Why Samsung Put Its HBM Base Die on a 4nm Logic Node
At Hot Chips 2026 last week, Samsung and SK hynix presented on the same day and attacked the same problem from opposite ends. Samsung's talk was about the base die. SK hynix's was about packaging.
Both start from one question: as stacks get taller, where does the heat go? Samsung's answer is the more structural one, because it changes what the bottom chip in an HBM stack actually is.
KEY TAKEAWAYS
1. Samsung's HBM4E is sampling at 14 Gbps per pin with 16 Gbps as the target. At 2,048 pins that is 3.6 TB/s now and 4.1 TB/s at target — twice the JEDEC HBM4 baseline.
2. The base die has moved to a 4nm foundry logic process. It is becoming a compute chip, not a wiring layer, and Samsung laid out a three-phase plan to keep pushing in that direction.
3. The forcing function is thermal. Base die power density goes from 0.5 to over 2.0 W/mm² in one generation — 4x — which is why heat management is moving into chip design rather than packaging.
First, separate the stages
Coverage of these announcements tends to collapse three different events into one. Development complete, sample shipment and mass production are not the same thing, and the difference matters when you are trying to size 2027 volume.
| Product | Stage | Pin speed | Stack |
|---|---|---|---|
| HBM4 | Shipping (announced Feb 12, 2026) | 11.7 Gbps, up to 13 | 12-high, 24-36 GB |
| HBM4E | Sampling (announced May 29, 2026) | 14 Gbps, scalable to 16 | 12-high, 48 GB |
| HBM5 | Roadmap | Not disclosed | 60+ GB, 6+ TB/s |
HBM4E is not in mass production. Samsung's own language is that it plans to begin mass production "aligned with customer schedules." The 16 Gbps and 4.0 TB/s figures shown at GTC in March are targets; the stable speed of the May samples is 14 Gbps.
The bandwidth numbers reconcile if you do the arithmetic yourself. At 2,048 pins per stack, 11.7 Gbps gives 3.0 TB/s, 14 Gbps gives 3.6 TB/s and 16 Gbps gives 4.1 TB/s. The JEDEC HBM4 baseline of 8 Gbps gives 2.05 TB/s, so Samsung is targeting roughly double the standard.
Why the base die went to 4nm
Start with the structure as it is today. An HBM stack is a set of DRAM core dies sitting on a base die, and that stack sits beside the XPU on a silicon interposer. Data travels sideways, across interposer traces, to get between them.
In that arrangement the base die was close to a wiring layer — it gathered signals and handed them to the XPU through the PHY — so a trailing logic node was good enough.
That changed with HBM4. Core dies use 1c DRAM, and the base die is built on a 4nm foundry process. At Hot Chips, Samsung explained the reasoning as a three-phase progression.
Phase 1 is area reclamation. Replacing the conventional PHY with an advanced-node D2D interface shrinks it substantially, and the memory controller moves down onto the base die as well. Samsung puts the result at 5-10% of XPU area recovered, which it translates into a 10-20% performance gain.
Phase 2 adds function — what Samsung calls aHBM. Temperature, voltage, process and aging sensors, on-die test circuitry, and a PHY and controller for external memory expansion all move into the base die, along with some processing elements. The justification offered was that context windows are growing 30x a year. When KV cache explodes, the first thing to cut is traffic between the XPU and memory.
The real driver is heat
One number in the deck explains the rest of it. Base die power density goes from 0.5 W/mm² at sHBM4E to over 2.0 W/mm² at sHBM5. Four times, in a single generation.
DRAM is temperature sensitive. Hotter cells need shorter refresh intervals, which costs both effective bandwidth and power. And the taller the stack, the fewer paths the lower dies have to get heat out. That is the structural weakness of the whole HBM concept.
Samsung's answer is a Heat Path Block built into the base die. Cover more than 50% of the PHY area with it, the company says, and peak temperature drops by over 35%. The notable part is not the number but the location: thermal management has moved from the package into the chip design flow.
zHBM: deleting the interposer
Phase 3 is zHBM — the right-hand panel of the cross-section above. The 2.5D interposer goes away and core dies stack vertically on the XPU itself, using wafer-on-wafer bonding and hybrid copper bonding.
Removing the trip across the interposer removes the serialization and data-alignment overhead that went with it, which is how I/O energy gets down to around 0.5 pJ/bit. Samsung's claimed figures against standard HBM4E are 230% more DRAM bandwidth and 70% better power efficiency. In a SiP with four zHBM stacks and a 1,200W GPU, that is roughly 100W saved — power that goes back to the XPU.
No commercialization date was given. And watch the baselines: at FMS 2026 on August 5, Samsung described zHBM as delivering up to 8x the performance of an HBM5-equipped accelerator. That is a different comparison from the 230% figure quoted against standard HBM4E. The two numbers do not belong in the same table.
SK hynix solves it in the package
| Samsung | SK hynix | |
|---|---|---|
| Topic | Base die, chip design | Advanced packaging, bonding |
| Thermal fix | Heat Path Block, peak temp down over 35% | i-HBM, thermal resistance down over 30% |
| Taller stacks | 16+ layers via hybrid bonding | 20+ layers, bump pitch under 18 um |
| Emphasis | Move compute into memory | Stress profiles of CoWoS-S, CoWoS-L, EMIB |
SK hynix's point was that the packaging choice determines the mechanical and thermal stress the HBM cube has to absorb — a genuinely different framing of the same constraint. It put HBM4 at 12-high in production with 16-high in qualification.
What I actually watch
| Signal | Why it matters |
|---|---|
| Foundry volume for base dies | Every HBM stack now carries a 4nm logic wafer with it |
| When HBM4E converts to mass production | "Customer schedules" is the swing factor for 2027 volume |
| Hybrid bonding tool orders | Both 16+ layers and zHBM depend on it; orders move before roadmaps do |
| How early thermal specs appear in decks | The further forward they move, the more binding the constraint is |
Risks to this view
- zHBM and aHBM are conference material, not products. No commercialization date was given for either.
- The 230% and 70% figures sit on a baseline the company chose. Change the baseline and the numbers change.
- I worked from published session coverage rather than the slides themselves.
- Memory roadmaps slip. Hybrid bonding alone has moved several times.
The center of gravity moved
HBM competition used to be about layer count and pin speed. This year's talks ask a different question: what goes into the chip at the bottom of the stack? Once memory vendors started buying leading-edge foundry capacity for it, half the answer was already settled.
Sources: Samsung Newsroom, HBM4 mass production shipment (Feb 12, 2026) and industry-first HBM4E samples (May 29, 2026); Samsung Semiconductor, HBM4E at NVIDIA GTC 2026 (March 2026); Samsung at FMS 2026 (Aug 5, 2026); Hot Chips 2026 session coverage, ServeTheHome (Aug 23, 2026). Everything here is from public sources.





Comments
Post a Comment