Why Two DGX Sparks Don't Run Twice as Fast

Link two DGX Spark 64GB boxes with a single cable and you get twice the memory and twice the memory bandwidth. NVIDIA says you get up to 1.7x the performance. The missing 0.3x is not a marketing rounding error. It is the cost of moving data between two machines, and it is the clearest way to understand what actually limits local AI hardware.

KEY TAKEAWAYS

1. Two clustered 64GB units double capacity (64 GB to 128 GB) and memory bandwidth (273 GB/s to 546 GB/s combined). The 1.7x performance figure is the maximum from NVIDIA's own Qwen 3.8 27B test, with no test conditions disclosed.

2. The bottleneck moves to the wire. Each unit's memory runs at 273 GB/s; the ConnectX-7 link between them is 200 Gb/s, or about 25 GB/s, roughly one-eleventh as fast.

3. There are two ways to buy 128 GB: two 64GB units at $9,998 in list prices, or one 128GB unit at a press-reported $6,950. One buys a single memory pool; the other buys double the bandwidth plus a communication tax.

What NVIDIA said, and what it left out

According to NVIDIA's October 2 blog post, the DGX Spark 64GB goes on sale October 23 from Acer, ASUS, Dell, Gigabyte, HP and MSI, starting at $4,999. It keeps the same GB10 Grace Blackwell chip and software stack as the 128GB model. One unit supports models up to 100 billion parameters. Two units joined over their built-in ConnectX-7 ports with a QSFP cable pool memory to 128 GB, support up to 200 billion parameters and, in NVIDIA's words, deliver "twice the memory bandwidth and up to 1.7x the performance."

The 1.7x comes from one place: "NVIDIA's Qwen 3.8 27B test," described as "up to." The post does not say what precision the model ran at, how long the context was, how many requests were in flight, or whether 1.7x refers to single-stream speed or total throughput. Some coverage, Tech Critter's among them, repeated the multiple without naming the model. There is no independent benchmark yet.

Two further numbers circulating in the press are not in NVIDIA's post. Hardware Busters and Guru3D report the 128GB Founders Edition now lists at $6,950. Guru3D, citing IT Home, reports that roughly 8 GB of the 64GB model is reserved for the system, leaving about 56 GB for weights and KV cache. I treat both as reported, not confirmed.

Two DGX Spark units: capacity and bandwidth double, performance up to 1.7x
What doubles and what does not. The 1.7x is NVIDIA's own test maximum; conditions undisclosed.

Capacity is the part that really doubles

Capacity decides whether a model fits, and that depends on bytes per parameter. For a 27-billion-parameter model, my rough math (weights only, excluding KV cache) is about 54 GB in BF16, 27 GB in FP8 and 14 GB at 4 bits.

Qwen3.8-27B is a dense 27B model, and its published weights are BF16, according to the Qwen model card on Hugging Face. At BF16, a single 64GB unit has almost no headroom left for context. At 4 bits, it fits on one box with room to spare. Same model, same "1.7x," entirely different meaning depending on precision. The same logic applies to NVIDIA's "100 billion parameters in 64 GB": that works out at about 4 bits per parameter (100B x 0.5 bytes = 50 GB, my math), not at every precision.

Bandwidth sets the speed once the model fits

Generating each token means reading the weights from memory again, which is why decode speed tracks memory bandwidth. I covered that mechanism in "Prefill, Decode, KV Cache: Where Token Cost Is Decided," so here are just the numbers.

NVIDIA's spec page lists DGX Spark memory as 256-bit LPDDR5X at 273 GB/s, with no distinction between the 64GB and 128GB versions. A DGX B200 lists 64 TB/s of HBM3e bandwidth across eight GPUs, or 8 TB/s per GPU by my division. That is a gap of about 29x per accelerator.

NVIDIA's own October 2025 developer blog shows what that means. Running Qwen3 235B in NVFP4 across two Sparks, prompt processing hit 23,477 tokens per second while generation ran at 11.73 tokens per second. The Register made a related point in January: NVIDIA's claimed 2.5x average software speedup since launch mostly helps prefill, while bandwidth-bound token generation barely moves.

The wire between two boxes

Clustering gives you 546 GB/s of combined memory bandwidth, but only if the model is split so each box reads half the weights. The moment you split it, the boxes must exchange intermediate results for every token. That traffic crosses a 200 Gb/s ConnectX-7 link, about 25 GB/s once converted to bytes.

Memory bandwidth inside one DGX Spark versus the link between two
On-board memory vs. the link between units, log scale. 25 GB/s is a unit conversion; B200 per-GPU figure is my division.

The vLLM documentation is explicit about the trade-off. Tensor parallelism, which splits work inside each layer, needs fast internode communication. Without NVLink, it recommends pipeline parallelism, which splits by layers, for lower communication overhead. Either way, synchronization is never free, and a large part of the gap between 2x and 1.7x plausibly sits there.

Illustrative calculation. Not actual company figures. Assume 27 GB of FP8 weights, read once per token, with bandwidth as the only limit. One unit: 27 GB / 273 GB/s = about 99 ms per token. Two units, perfect split: about 49 ms. If the real result is 1.7x, a token takes about 58 ms, implying roughly 9 ms per token of communication and synchronization.
Illustrative breakdown of per-token time for one and two DGX Spark units
Illustrative calculation. Not actual company figures.

NVIDIA has not said whether 1.7x is per-request latency or aggregate throughput across concurrent requests, so this is a way to size the problem, not a measurement.

Two ways to buy 128 GB

Two 64GB units cost $9,998 at NVIDIA's starting price. One 128GB unit costs a reported $6,950. That is about $78 per GB versus $54 per GB by my math. Back2Gaming notes the 128GB model launched at $3,999 in October 2025, and coverage ties the increases to tight memory supply.

Price of two 64GB DGX Spark units versus one 128GB unit
Two routes to 128 GB. 64GB price from NVIDIA; 128GB price as reported by the press.

The pricier route buys 546 GB/s of combined bandwidth and two sets of compute. The cheaper one buys a single 128 GB memory pool with no link to cross. Same capacity on paper, different machines underneath: one is a capacity purchase, the other a bandwidth purchase with a communication tax attached.

What I actually watch

CheckpointWhat would move the view
NVIDIA test disclosurePrecision, context length, concurrency behind 1.7x
Independent reviews after Oct 23One-unit vs. two-unit tokens per second
Official 128GB pricingWhether $6,950 appears on NVIDIA's own store
LPDDR5X pricingHow sensitive local AI boxes are to memory cost

Value chain read-through

LayerLink to two-unit scaling
LPDDR5XSource of capacity and bandwidth
ConnectX-7 NICThe path between units
QSFP cablingDirect unit-to-unit link
GB10 chipDouble compute, plus sync cost

Risks to this view

• The 1.7x is NVIDIA's own test maximum, with undisclosed conditions and no independent verification.
• The $6,950 price and the roughly 56 GB usable memory are press reports not confirmed in NVIDIA's primary materials.
• The 99 ms / 49 ms / 9 ms breakdown is an illustrative calculation, not actual company figures.
• Software updates can change performance on the same hardware, so today's multiples are not fixed.

Next, the other end of the memory spectrum: if on-chip SRAM is that fast, does it make HBM unnecessary? Speed and capacity turn out to be different problems.

Sources: NVIDIA blog, "NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI" (Oct 2, 2026); NVIDIA DGX Spark and DGX B200 spec pages (accessed Oct 4, 2026); NVIDIA Developer Blog, "How NVIDIA DGX Spark's Performance Enables Intensive AI Tasks" (Oct 24, 2025); Qwen3.8-27B model card, Hugging Face; vLLM docs, Parallelism and Scaling; Hardware Busters, Guru3D, Back2Gaming (Oct 2, 2026); Tech Critter (Oct 3, 2026); The Register (Jan 5, 2026).

Disclaimer: This post is for informational and educational purposes only. It does not constitute investment advice or a recommendation to buy or sell any security. All investment decisions are your own responsibility.

Comments

Popular posts from this blog

Why Nvidia's Inference GPU Skips HBM for GDDR7

Korea's August Chip Exports Hit a Record $46.7B. Volume Moved Too

DDR4 Costs More Than DDR5 — Unless You're Actually Buying It