Jalapeno's HBM4 Runs at 10 Gbps. Samsung Rates 11.7.
OpenAI published the first benchmarks for its in-house inference chip, Jalapeño, at Hot Chips on August 25. Most of the coverage led with the supplier question, since SemiAnalysis said the HBM4 is likely Samsung's. The more useful number was sitting in the spec sheet the whole time: divide 15.4 TB/s by six stacks and you get an operating point that sits below what Samsung publicly rates its own HBM4 at.
KEY TAKEAWAYS
1. Jalapeño pairs one compute die with six HBM4 stacks for 216 GiB and 15.4 TB/s. That is 2.57 TB/s per stack, which back-solves to roughly 10 Gbps per pin across a 2,048-bit interface.
2. Samsung announced 11.7 Gbps and 3.3 TB/s per stack when it started shipping HBM4 in February 2026. NVIDIA publishes 10.8 Gbps for Rubin. If the supplier attribution holds, the part is running below its rated ceiling.
3. The performance claim swings from 1.5x to 104.3x inside the same deck depending on operating point and power basis. The condition matters more than the multiple.
Two sources, two confidence levels
The news came from two places. Richard Ho, OpenAI's hardware lead, gave Bloomberg an interview. Separately, OpenAI presented at Hot Chips 2026 and SemiAnalysis published an architectural breakdown after running its InferenceX suite in OpenAI's lab. Almost everything downstream is a rewrite of those two.
What OpenAI stated directly: Jalapeño goes into limited production serving of its own models starting later this year, it was co-developed with Broadcom, it is inference-only, it is rated at 700W, and it beat NVIDIA's GB300 on throughput per watt and end-to-end latency.
What is inference by analysts rather than company statement: the six-stack, 216 GiB, 15.4 TB/s memory configuration, the TSMC 3nm-class process, and the identification of Samsung as the HBM supplier. SemiAnalysis used the word "likely." Neither OpenAI nor Samsung has confirmed it.
Divide 15.4 by six
HBM4 under JEDEC's JESD238 runs a 2,048-bit interface per stack at 8.0 Gbps per pin, which yields 2.0 TB/s per stack. That is the floor the standard defines, not the ceiling vendors ship.
Jalapeño's 15.4 TB/s across six stacks is 2.57 TB/s each. Run the interface width backwards and the pin rate lands at about 10 Gbps, roughly 25% above the JEDEC baseline. Capacity divides the same way: 216 GiB over six stacks is 36 GB per stack, which is a 12-high configuration.
None of this requires inside knowledge. It is arithmetic on two published numbers, and it takes about a minute.
Where Samsung's 11.7 Gbps went
When Samsung announced HBM4 mass-production shipments in February 2026, it published 11.7 Gbps per pin, 3.3 TB/s per stack, and 36 GB at 12-high, with headroom to 13 Gbps. NVIDIA's published Rubin spec is 10.8 Gbps, which is how 8 stacks reach 22 TB/s per package.
The back-solved Jalapeño number, 10 Gbps, is below both. That is not a contradiction, and it is the part worth internalizing: HBM runs at the operating point the host chip's design budget allows, not at the vendor's headline rating. Power envelope, thermals, and package signal integrity set that point. Jalapeño was designed around a 700W rating, and the memory configuration follows from that constraint rather than from what the DRAM can do in isolation.
The read-through for anyone tracking the memory makers: headline spec wins do not translate into revenue on their own. Revenue is stacks per package multiplied by packages shipped, at whatever speed grade the customer actually orders. A vendor can hold the fastest published number and still ship fewer stacks than a competitor.
1.5x, 1.9x, and 104.3x are all in the same deck
The headline figures are 1.5x to 1.9x more throughput per kilowatt and 1.7x to 3.6x lower end-to-end latency against GB200 and GB300 rack systems. At low-latency operating points, where the GB300 is pushed to its fastest time-between-tokens setting, OpenAI reports 8.6x to 104.3x. When the GB300 runs multi-token prediction, which is standard in production deployments, the lead compresses to about 1.5x.
The power basis moves too. Headline numbers normalize to rated package TDP, 700W against 1,400W. An appendix comparison using all-in utility power per accelerator puts Jalapeño at 1.18 kW and the GB300 at 2.55 kW. OpenAI separately states that Jalapeño's measured sustained power stayed at or below 550W.
Test models were GPT-OSS 120B, DeepSeek R1 670B, and Moonshot's 1-trillion-parameter Kimi K2.5. No frontier model was included, and Vera Rubin, the same-generation HBM4 platform, was not tested.
One more claimant on HBM4
For the memory cycle, the leaderboard is not the story. A new buyer is.
OpenAI signed a 10 GW deployment agreement with Broadcom last October. To the extent that converts into orders, HBM4 demand grows. A framework agreement is not booked volume, so the size is open.
Supply is not. Samsung, SK hynix and Micron have effectively sold their HBM capacity through 2027. At the same Hot Chips conference on August 23, Micron said HBM consumes roughly three times the wafer area of DDR5 for equivalent capacity, and that penalty widens each generation. When a new claimant arrives against fixed capacity, pricing power sits with the suppliers.
That is the frame I would use here, rather than treating it as a single-company order announcement.
What I actually watch
| Item | Why it matters | When |
|---|---|---|
| Supplier confirmation | Currently an analyst inference, not a disclosure | Either company's statement |
| Move from engineering samples | Nothing ships until this happens | 2027 ramp |
| Second-generation tapeout | Sets volume and HBM spec for the real deployment | Expected within months |
| A Vera Rubin comparison | The only like-for-like HBM4 matchup | Next benchmark release |
Value chain read-through
| Segment | Effect | Indicator to track |
|---|---|---|
| HBM supply | New claimant, capacity sold through 2027 | Capacity and pre-sale commentary from the three vendors |
| Foundry | More contention for TSMC 3nm-class wafers | TSMC advanced-node utilization |
| Advanced packaging | Additional CoWoS-class demand | TSMC packaging expansion plans |
| Power and cooling | 1.18 kW all-in per accelerator | Data center power contracts |
Risks to this view
— The supplier is unconfirmed. If the attribution is wrong, the entire framing above changes.
— OpenAI supplied most of the performance data. SemiAnalysis verified some runs on site, not all of them.
— This is inference only. Training remains uncontested, and the same-generation Rubin comparison has not been run.
— Jalapeño is reported to be at engineering-sample stage with a 2027 production ramp.
— OpenAI explicitly said it still needs NVIDIA, and on August 17 agreed to take up to $105 billion in NVIDIA financing for an Ohio data center campus. Vendor diversification and vendor dependence are running at the same time.
Closing
What the deck establishes is that a first-generation ASIC reached a usable inference operating point in about sixteen months. What it does not establish is who supplies the memory, or at what volume.
The bigger the headline number, the more useful it is to check the denominator. Next post: what "HBM capacity sold out through 2027" actually means in wafer-area terms.
Sources: OpenAI Hot Chips 2026 presentation; Bloomberg interview (Aug 25, 2026); SemiAnalysis; Tom's Hardware (Aug 25, 2026); JEDEC JESD238; Samsung announcement (Feb 2026); NVIDIA published specifications; TrendForce. Everything here is from public sources.
Disclaimer: This post is for informational and educational purposes only. It does not constitute investment advice or a recommendation to buy or sell any security. All investment decisions are your own responsibility.




Comments
Post a Comment