Why Two DGX Sparks Don't Run Twice as Fast
Link two DGX Spark 64GB boxes with a single cable and you get twice the memory and twice the memory bandwidth. NVIDIA says you get up to 1.7x the performance. The missing 0.3x is not a marketing rounding error. It is the cost of moving data between two machines, and it is the clearest way to understand what actually limits local AI hardware.
KEY TAKEAWAYS
1. Two clustered 64GB units double capacity (64 GB to 128 GB) and memory bandwidth (273 GB/s to 546 GB/s combined). The 1.7x performance figure is the maximum from NVIDIA's own Qwen 3.8 27B test, with no test conditions disclosed.
2. The bottleneck moves to the wire. Each unit's memory runs at 273 GB/s; the ConnectX-7 link between them is 200 Gb/s, or about 25 GB/s, roughly one-eleventh as fast.
3. There are two ways to buy 128 GB: two 64GB units at $9,998 in list prices, or one 128GB unit at a press-reported $6,950. One buys a single memory pool; the other buys double the bandwidth plus a communication tax.
What NVIDIA said, and what it left out
According to NVIDIA's October 2 blog post, the DGX Spark 64GB goes on sale October 23 from Acer, ASUS, Dell, Gigabyte, HP and MSI, starting at $4,999. It keeps the same GB10 Grace Blackwell chip and software stack as the 128GB model. One unit supports models up to 100 billion parameters. Two units joined over their built-in ConnectX-7 ports with a QSFP cable pool memory to 128 GB, support up to 200 billion parameters and, in NVIDIA's words, deliver "twice the memory bandwidth and up to 1.7x the performance."
The 1.7x comes from one place: "NVIDIA's Qwen 3.8 27B test," described as "up to." The post does not say what precision the model ran at, how long the context was, how many requests were in flight, or whether 1.7x refers to single-stream speed or total throughput. Some coverage, Tech Critter's among them, repeated the multiple without naming the model. There is no independent benchmark yet.
Two further numbers circulating in the press are not in NVIDIA's post. Hardware Busters and Guru3D report the 128GB Founders Edition now lists at $6,950. Guru3D, citing IT Home, reports that roughly 8 GB of the 64GB model is reserved for the system, leaving about 56 GB for weights and KV cache. I treat both as reported, not confirmed.

Capacity is the part that really doubles
Capacity decides whether a model fits, and that depends on bytes per parameter. For a 27-billion-parameter model, my rough math (weights only, excluding KV cache) is about 54 GB in BF16, 27 GB in FP8 and 14 GB at 4 bits.
Qwen3.8-27B is a dense 27B model, and its published weights are BF16, according to the Qwen model card on Hugging Face. At BF16, a single 64GB unit has almost no headroom left for context. At 4 bits, it fits on one box with room to spare. Same model, same "1.7x," entirely different meaning depending on precision. The same logic applies to NVIDIA's "100 billion parameters in 64 GB": that works out at about 4 bits per parameter (100B x 0.5 bytes = 50 GB, my math), not at every precision.
Bandwidth sets the speed once the model fits
Generating each token means reading the weights from memory again, which is why decode speed tracks memory bandwidth. I covered that mechanism in "Prefill, Decode, KV Cache: Where Token Cost Is Decided," so here are just the numbers.
NVIDIA's spec page lists DGX Spark memory as 256-bit LPDDR5X at 273 GB/s, with no distinction between the 64GB and 128GB versions. A DGX B200 lists 64 TB/s of HBM3e bandwidth across eight GPUs, or 8 TB/s per GPU by my division. That is a gap of about 29x per accelerator.
NVIDIA's own October 2025 developer blog shows what that means. Running Qwen3 235B in NVFP4 across two Sparks, prompt processing hit 23,477 tokens per second while generation ran at 11.73 tokens per second. The Register made a related point in January: NVIDIA's claimed 2.5x average software speedup since launch mostly helps prefill, while bandwidth-bound token generation barely moves.
The wire between two boxes
Clustering gives you 546 GB/s of combined memory bandwidth, but only if the model is split so each box reads half the weights. The moment you split it, the boxes must exchange intermediate results for every token. That traffic crosses a 200 Gb/s ConnectX-7 link, about 25 GB/s once converted to bytes.

The vLLM documentation is explicit about the trade-off. Tensor parallelism, which splits work inside each layer, needs fast internode communication. Without NVLink, it recommends pipeline parallelism, which splits by layers, for lower communication overhead. Either way, synchronization is never free, and a large part of the gap between 2x and 1.7x plausibly sits there.

NVIDIA has not said whether 1.7x is per-request latency or aggregate throughput across concurrent requests, so this is a way to size the problem, not a measurement.
Two ways to buy 128 GB
Two 64GB units cost $9,998 at NVIDIA's starting price. One 128GB unit costs a reported $6,950. That is about $78 per GB versus $54 per GB by my math. Back2Gaming notes the 128GB model launched at $3,999 in October 2025, and coverage ties the increases to tight memory supply.

The pricier route buys 546 GB/s of combined bandwidth and two sets of compute. The cheaper one buys a single 128 GB memory pool with no link to cross. Same capacity on paper, different machines underneath: one is a capacity purchase, the other a bandwidth purchase with a communication tax attached.
What I actually watch
| Checkpoint | What would move the view |
|---|---|
| NVIDIA test disclosure | Precision, context length, concurrency behind 1.7x |
| Independent reviews after Oct 23 | One-unit vs. two-unit tokens per second |
| Official 128GB pricing | Whether $6,950 appears on NVIDIA's own store |
| LPDDR5X pricing | How sensitive local AI boxes are to memory cost |
Value chain read-through
| Layer | Link to two-unit scaling |
|---|---|
| LPDDR5X | Source of capacity and bandwidth |
| ConnectX-7 NIC | The path between units |
| QSFP cabling | Direct unit-to-unit link |
| GB10 chip | Double compute, plus sync cost |
Risks to this view
• The 1.7x is NVIDIA's own test maximum, with undisclosed conditions and no independent verification.
• The $6,950 price and the roughly 56 GB usable memory are press reports not confirmed in NVIDIA's primary materials.
• The 99 ms / 49 ms / 9 ms breakdown is an illustrative calculation, not actual company figures.
• Software updates can change performance on the same hardware, so today's multiples are not fixed.
Next, the other end of the memory spectrum: if on-chip SRAM is that fast, does it make HBM unnecessary? Speed and capacity turn out to be different problems.
Sources: NVIDIA blog, "NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI" (Oct 2, 2026); NVIDIA DGX Spark and DGX B200 spec pages (accessed Oct 4, 2026); NVIDIA Developer Blog, "How NVIDIA DGX Spark's Performance Enables Intensive AI Tasks" (Oct 24, 2025); Qwen3.8-27B model card, Hugging Face; vLLM docs, Parallelism and Scaling; Hardware Busters, Guru3D, Back2Gaming (Oct 2, 2026); Tech Critter (Oct 3, 2026); The Register (Jan 5, 2026).
Comments
Post a Comment