Nebius Buys Inferize: The Cost of an 84-Second Cold Start
The AI capex debate is usually about how many GPUs get bought. The quieter question is how many hours those GPUs actually serve traffic. On October 1, Nebius bought a startup whose entire pitch is shrinking the dead time between "a request arrives" and "the model is ready to answer." That dead time is the cold start, and it is a bigger cost line than it looks.
KEY TAKEAWAYS
1. Loading a model is slow relative to serving it. In an OSDI 2024 experiment, loading LLaMA-2-70B onto 8 GPUs took 84 seconds; generating one token typically takes under 0.1 seconds.
2. Slow loading forces operators to keep GPUs warm for peak demand. In a simple illustrative model, scaling with demand serves the same day with 32% fewer GPU-hours.
3. Nebius spent about $5.7bn on capex in Q2 2026, roughly 10x its $0.58bn quarterly revenue. Inferize is its third inference-stack addition this year, aimed at getting more work out of that hardware.
What Nebius actually announced
According to Nebius's October 1 press release, Inferize's technology and team have joined Nebius Token Factory, the company's managed inference platform. Terms were not disclosed. Nebius describes the problem as cold starts, the time a model needs to load before it can serve requests, and calls the resulting idle hardware an "idle GPU tax."
The deal follows two earlier moves. On May 1, Nebius agreed to acquire Eigen AI, a model and inference optimization company, for an aggregate value of about $643 million at signing in cash and stock. On May 12, Clarifai's core team joined Nebius under a license to Clarifai's inference and compute orchestration technology, with commercial terms undisclosed.
What the release does not contain is a number. There is no claimed speedup, no cost saving, and no third-party benchmark. Inferize was founded in January 2026 and had a working prototype within three months. So the deal tells us where Nebius thinks the bottleneck is, not yet how much it can be moved.
How long is a cold start?
A large language model is, physically, a set of weight files measured in tens to hundreds of gigabytes. Before it can answer anything, those weights have to be read from storage, copied into GPU memory, and initialized.
A useful public measurement comes from the ServerlessLLM paper (Fu et al., USENIX OSDI 2024). Even with the checkpoint already sitting on local NVMe, default PyTorch loading took 34 seconds for OPT-30B on 4 GPUs and 84 seconds for LLaMA-2-70B on 8 GPUs. The same paper notes that generating a token usually takes under 100 milliseconds.

That is three orders of magnitude between getting ready and doing the work. The paper also notes that several serverless providers, Bloomberg among them, have publicly described LLM initialization latencies of tens of seconds.
It is a memory hierarchy problem
The biggest variable is where the weights live when the request arrives. Take the paper's example of a 130 GB checkpoint (LLaMA-2-70B in FP16) and divide it by the bandwidth of each tier the paper cites:
• Remote object storage at 5 GB/s: at least 26 seconds (the paper's figure).
• Local NVMe in RAID 0 at about 60 GB/s: about 2.2 seconds (my division).
• Host DRAM to 8 GPUs over PCIe 5.0, 512 GB/s aggregate: about 0.25 seconds (my division).

These are floors, not forecasts. Memory allocation and tensor setup add time on top. ServerlessLLM's answer was a loading-optimized checkpoint format plus using the DRAM and SSDs already inside GPU servers as a tiered cache. For the 70B model, it reported loading 8.2x faster than PyTorch. The takeaway for hardware readers: cold-start engineering is largely about placing bytes in the right tier of DRAM and flash before they are needed.
The idle GPU tax, illustrated
If a model takes minutes to bring up, an operator cannot wait for demand to arrive. It keeps capacity warm. The ServerlessLLM authors list over-subscribing GPUs as the prevailing workaround. To show the scale of that choice, here is a deliberately simple model.
Option A burns 2,640 GPU-hours at 59% utilization. Option B burns 1,806 GPU-hours at 86% utilization. That is 32% fewer GPU-hours for identical demand. The catch is that Option B only works if new replicas come up well within the scaling window. That is the gap Inferize says it closes when it talks about making inference "elastic."

Why Nebius cares: 10x capex to revenue
Nebius's Q2 2026 shareholder letter (filed on Form 6-K on August 12) shows the economics. Group revenue was $582.3 million, up 454% year over year. Annualized run-rate revenue reached $3.0bn at the end of June. Capital expenditure was approximately $5.7bn, mainly GPUs. By my calculation, that is about 9.8x quarterly revenue.

The letter also says production inference workloads more than tripled in Q2. That mix shift matters. Training runs tend to keep GPUs busy for long stretches. Inference follows user traffic, which spikes and fades. As inference grows, returns depend less on how much hardware a cloud buys and more on how little of it sits idle.
What I actually watch
| Checkpoint | What would move the view |
|---|---|
| Nebius Q3 shareholder letter | Token Factory inference volume; any Inferize integration metric |
| Published cold-start times | Start-up time by model size, not just "faster" |
| Competing serverless GPU offers | Start-up latency and billing granularity |
| Inference server configs | More DRAM and NVMe per node for checkpoint caching |
Value chain read-through
| Layer | Link to cold starts |
|---|---|
| GPUs | Utilization, not just unit count |
| Server DRAM | Hot cache tier for weights |
| NVMe SSDs | Local checkpoint storage |
| Networking | Remote download bandwidth |
For readers who follow Korean memory makers through the customs data I track on this blog: the DRAM and NVMe rows are the ones to keep in mind. If elastic inference becomes the norm, more of the model-loading path runs through server memory and flash rather than idle HBM.
Risks to this view
• Inferize's impact is described only by Nebius. No savings figure or benchmark has been published.
• Higher utilization cuts both ways. It can also mean fewer GPUs for the same demand, so it is not automatically bullish for GPU volumes.
• This is Nebius's third inference-stack deal this year. Integrating three teams into one platform carries execution risk.
• Capex running far ahead of revenue makes the model sensitive to demand slowdowns and financing conditions.
The cold start sits at the front of the inference cost stack, before prefill, decode, and KV-cache management even begin. Next, I'll look at the power side of the same problem: data center developers ordering their own on-site generation.
Sources: Nebius press releases (Oct 1, May 12 and May 1, 2026); Nebius Q2 2026 shareholder letter, SEC Form 6-K Exhibit 99.2 (Aug 12, 2026); Fu et al., "ServerlessLLM: Low-Latency Serverless Inference for Large Language Models," USENIX OSDI 2024.
Comments
Post a Comment