9,600 Chips, One 2PB Pool: What Hot Chips 2026 Day 2 Was About

 Hot Chips 2026 wrapped up at Stanford on August 25, and the second conference day carried one message across three sessions: the binding constraint in AI systems is shifting from how fast a chip computes to how efficiently data moves between chips. Networking, hyperscaler custom silicon and memory all told the same story from different angles.

KEY TAKEAWAYS

1. Broadcom's Thor Ultra NIC filled 98.9% of an 800G link in measured TCP throughput, and NVIDIA said its co-packaged optics are in production. The networking session drew the biggest strategic claims of the event.

2. Google's TPU 8t superpod pools 2PB of HBM across 9,600 chips as one shared memory space. Meta's MTIA roadmap triples HBM bandwidth in three generations, from 9.2 to 27.6 TB/s. Custom silicon is increasingly a data-movement strategy, not just a cost play.

3. Samsung productized the first LPDDR-based processing-in-memory part, tripling token generation from 27 to 81.3 tokens per second in a Llama 3.1 demo. Korean startup XCENA put 3,072 RISC-V cores inside a CXL memory device.



Why day two mattered

Day one belonged to compute: NVIDIA Rubin, AMD MI400, Intel Crescent Island. Day two grouped networking, the hyperscaler accelerator programs and memory into back-to-back sessions, and the framing converged. As models scale to trillion-parameter, long-context, agentic workloads, the time budget moves into decode and data movement rather than raw FLOPS.

Two clarifications before the numbers. First, NVIDIA's Groq 3 LPX production announcement on day one produced two figures that got mixed up in coverage: 3,400 tokens per second is a measured per-request output rate on Gemma 4 31B with a 100,000-token context (Artificial Analysis), while 11,000 tokens per second is the aggregate decode throughput of one rack, from the presentation slides. Different units, both real. Second, the "NVIDIA challenger" framing around TPU and MTIA misses what the talks actually emphasized: hyperscalers are optimizing their own workload's system-level bottlenecks, and those bottlenecks are memory and interconnect.

Networking: 98.9% of the link, and optics in production

Broadcom's Thor Ultra pairs a PCIe Gen6 x16 host interface with eight 100G SerDes for an 800G-class Ethernet NIC. The measured numbers stood out: 791 Gbps unidirectional TCP, or 98.9% of the 800G link, and 1,558 Gbps bidirectional RDMA, about 97.6% of the 1.6 Tb/s aggregate. The gap between link spec and delivered throughput has nearly closed.

NVIDIA's BlueField-4 is an 800 Gb/s-class DPU built around a 64-core Grace CPU, demonstrated running as a 7 Tb/s platform-level DPU inside a Vera Rubin system — a platform NVIDIA describes as seven chips across five racks. The Spectrum-X talk that followed argued an AI factory needs five purpose-built networks (scale-up, scale-out, scale-across, scale-in and an AI context tier), and mentioned in passing that NVIDIA's co-packaged optics solutions are in production. That one line is worth tracking: CPO moving from roadmap to production reshapes the optical module value chain.

Hyperscaler silicon: the shared-memory-pool race

Google's TPU 8t, the training half of its eighth-generation family, is best read as a memory architecture. One superpod holds 9,600 chips sharing 2PB of HBM as a single pool, delivering 121 EFLOPS of FP4 compute per Google's figures. Each chip carries 216GB of HBM at 6.5 TB/s. Against Ironwood's 9,216 chips and 1.77PB, chip count grew just 4% while pod compute roughly tripled — the leverage is in the interconnect and the shared pool, not the chip count. Google says TPU performance has improved a million-fold over the program's history. The inference sibling, TPU 8i, pairs two-to-one with Google's own Axion CPUs.



Meta's MTIA 400 delivers 6 PFLOPS of FP8 — five times the prior generation — with eight stacks of HBM3E totaling 288GB (1.3x) at 9.4 TB/s per the Hot Chips presentation, in a 1,200W module. Seventy-two devices form one rack-scale scale-up domain. The roadmap is the tell: MTIA 450 (early 2027) doubles bandwidth to 18.4 TB/s, and MTIA 500 (2027) reaches 27.6 TB/s with 384–512GB of capacity. Bandwidth and capacity scale ahead of FLOPS, because transformer decode is bandwidth-bound, not compute-bound.

Memory that computes: Samsung PIM and XCENA's CXL device

Samsung opened its talk with a cost chart: HBM's share of AI chip component cost rose from 52% in Q1 2024 to 63% by Q4 2025. Bandwidth is necessary; HBM prices are the problem. Its answer is LPDDR5X-PIM — sixteen MAC-equipped PIM blocks inside the DRAM banks, so data gets computed where it sits instead of crossing to the processor. Internal PIM bandwidth reaches 614 GB/s, eight times standard LPDDR5X. In a Llama 3.1 demo, token generation rose from 27 to 81.3 tokens per second and task completion time fell from 12.3 to 5.4 seconds. It ships in the same JEDEC 561-ball package as standard LPDDR5X, so it drops in without a system redesign. Samsung calls it the first LPDDR-based PIM product, per its own announcement.


In the same session, Korean startup XCENA presented jointly with Samsung on MX1, a CXL Type 3 computational memory device built on Samsung Foundry's 4nm process. It combines memory expansion over four DDR5-8400 channels up to 2TB, SSD-backed capacity tiering and near-memory processing across 3,072 RISC-V cores. Across six data-processing kernels — compression, decompression, Parquet decoding and filtering among them — XCENA reported up to 4.7x higher throughput and 18.7x better energy efficiency than a host CPU pulling the same data over CXL.

These three threads — better networks, bigger shared pools, compute inside memory — answer the same question. A byte that stays local is a byte that never touches a congested interconnect. That is why memory innovation now reads as an extension of networking innovation.

What I actually watch

CheckpointWhat to look for
CPO volume rampNetworking revenue growth on NVIDIA's earnings calls; optical component order commentary
TPU 8t availabilityGeneral availability in 2H 2026 and HBM procurement scale — 2PB per pod is supportive for total HBM demand
LPDDR5X-PIM adoptionNamed customers or sockets (server NPUs, on-device). Today this is a productization announcement, not a design win
MTIA 450 deploymentWhether the early-2027 schedule holds, and who supplies the HBM

Value chain read-through

SegmentImplication
Networking (NICs, switches, CPO)Value spreads from single accelerators to the whole data-movement stack
HBM suppliersBigger shared pools and bandwidth-first roadmaps like MTIA's are supportive for aggregate HBM demand
PIM and CXLSamsung takes the first LPDDR-PIM product to market; a Korean fabless (XCENA) shows up in the CXL ecosystem on Samsung Foundry silicon

Risks to this view

— Announcements are not adoption. LPDDR5X-PIM has no disclosed customers or volumes, and MX1 is early-stage.

— TPU 8t and MTIA figures are company-stated, not third-party measured, and custom-silicon roadmaps slip often.

— If PIM and CXL scale, the long-run effect is less dependence on expensive HBM per unit of work. Near-term HBM demand and long-term mix are two different questions, and this post is bullish only on the first.

The through-line from day two is simple: the next round goes not to whoever builds the fastest chip, but to whoever moves the least data — and moves it best. Next up, I will look at what "co-packaged optics in production" means for the optical module value chain.

Sources: Samsung, Google, Meta, NVIDIA, Broadcom and XCENA presentations and press materials; ServeTheHome and Tom's Hardware Hot Chips 2026 session coverage; Meta engineering blog (March 2026); Google Cloud Next 2026 announcements.

Disclaimer: This post is for informational and educational purposes only. It does not constitute investment advice or a recommendation to buy or sell any security. All investment decisions are your own responsibility.

Comments

Popular posts from this blog

Why Nvidia's Inference GPU Skips HBM for GDDR7

Korea's August Chip Exports Hit a Record $46.7B. Volume Moved Too

DDR4 Costs More Than DDR5 — Unless You're Actually Buying It