M5 Ultra vs Dual DGX Spark: Which Wins at Local LLMs?
M5 Ultra Mac Studio vs clustered dual NVIDIA DGX Spark, benchmarked on DeepSeek V4 Flash and Qwen for real local LLM speed.

The short answer
Neither machine wins outright. In head-to-head testing, the M5 Ultra Mac Studio (256GB unified memory) and a cluster of two NVIDIA DGX Sparks (128GB each, linked over a 200GB RDMA cable) produced nearly identical token generation speeds on large local models like DeepSeek V4 Flash, both landing around 34 to 38 tokens per second. But the Sparks crushed the Mac on prompt processing, reading long inputs two to three times faster, while the Mac held an edge in raw output speed on some models and in multi-user efficiency up to a point. Which one is “better” depends entirely on whether your workload is read-heavy or write-heavy.
TL;DR
- Token generation speed was close to a tie, with the M5 Ultra and dual DGX Spark cluster both producing roughly 34 to 38 tokens per second on DeepSeek V4 Flash in single-user tests.
- Prompt processing is where the Sparks pull ahead decisively, reading a 32,000-token codebase in about 17 seconds versus roughly 44 to 50 seconds on the Mac, a gap that widens further at longer context lengths.
- Memory bandwidth explains most of the generation-speed story, since Apple rates the M5 Ultra at 1.2 terabytes per second versus about 273 gigabytes per second per Spark, even though the Sparks partially offset that by splitting work across two GPUs.
- Software choice changes the results substantially, since switching DeepSeek from llama.cpp to MLX on the Mac boosted output from about 40 to 53 tokens per second, a 34% jump from the engine alone.
- Prompt caching erases most of the real-world pain, because once a codebase or long context is cached, follow-up questions only add new tokens rather than re-reading everything, shrinking the Mac’s disadvantage dramatically in normal coding-agent use.
- Multi-user throughput favors the Sparks at scale, with the Mac’s generation speed peaking around four concurrent users (66 tokens/sec total) before dropping to 46 at eight, while the Spark cluster kept climbing past that point.
- Price and footprint are close enough to matter, with the tested Mac configuration running about $14,000 (including an 8TB drive) against roughly $10,000 for two Sparks plus the cost of the RDMA cable linking them.
Other agents ship a demo. Remy ships an app.
Real backend. Real database. Real auth. Real plumbing. Remy has it all.
How does a dual DGX Spark cluster actually work?
Two DGX Sparks don’t merge into one computer automatically. They need a specific software setup, and in this case that was vLLM, NVIDIA’s go-to inference engine for clustered GPU setups (llama.cpp also runs on Sparks, but vLLM is the more common choice for this kind of split). The two units communicate over a direct cable connection running RoCE (RDMA over Converged Ethernet), which lets them write directly into each other’s memory rather than routing traffic through a slower network stack. The cable tested was rated for 200GB but measured closer to 111GB in practice.
The specific technique used is called tensor parallelism. Every layer of the model gets split in half, with one half processed on each Spark. For every single token generated, both machines compute their half of the layer, then swap results over the cable before the next layer can start. That means constant back-and-forth communication for every token, which is why cable speed and latency matter so much for this architecture.
This split matters because modern large models often don’t fit on a single 128GB Spark. DeepSeek V4 Flash runs around 150 to 160GB, too large for one unit once you account for the extra memory needed for context. Spread across two Sparks, each node ends up holding a bit over 100GB, leaving room to operate. The Mac, by contrast, just loads the whole quantized model into its single pool of unified memory.
Why does the Mac lose so badly on prompt processing?
Inference happens in two distinct phases, and they’re bottlenecked by completely different hardware characteristics. The first phase, prompt processing (also called “prefill”), is pure matrix multiplication: the model reads your entire prompt at once, and that’s a compute-bound task where raw GPU horsepower wins. The second phase, token generation, means writing the answer one token at a time, and each token requires another pass through the model’s memory, making that phase bound by memory bandwidth rather than compute.
The DGX Sparks, with two GPUs working in parallel, have more raw compute to throw at reading long prompts. In testing, a 32,000-token codebase (14 Python files) took about 17 seconds to process on the dual Spark cluster versus around 44 to 50 seconds on the M5 Ultra, a gap of roughly 2.4x to 3x in the Sparks’ favor. Push the prompt to 128,000 tokens and the Sparks finished in about 72 seconds, while the Mac’s pace at smaller sizes suggested it would have taken more than three minutes at that size (it wasn’t tested at full length because the pattern was already clear).
That compute gap doesn’t show up the same way on short prompts. At 2,000 tokens, both machines finish fast enough that the difference (under two seconds either way) is imperceptible in normal use.
Does memory bandwidth explain the generation speed results?
Largely, yes. Apple specifies the M5 Ultra’s unified memory bandwidth at 1.2 terabytes per second. Each DGX Spark is rated at about 273 gigabytes per second, and even combined across two units with a fast interconnect, that’s a meaningful gap in raw memory throughput, the resource that governs token generation speed. That bandwidth advantage is likely why the Mac matched or beat the dual-Spark cluster on writing speed despite the Sparks having more total compute.
But the comparison isn’t purely about hardware specs, because the software engine running on top makes a real difference. On the Mac, DeepSeek V4 Flash produced about 40 tokens per second under llama.cpp but jumped to 53 tokens per second under MLX, Apple’s own machine learning framework, a 34% improvement just from switching software. When llama.cpp was used on both machines under equivalent serving conditions, the Mac and the Spark cluster landed within 2% of each other on generation speed. On Qwen, the Mac’s 45 tokens per second clearly beat the Sparks’ 38, a result attributed mainly to memory bandwidth.
Does prompt caching change the real-world picture?
Yes, substantially. The long prefill times quoted above are worst-case, cold-start numbers: what happens when a session begins and the entire prompt has never been seen before. In actual use, especially with coding agents, a session’s context gets cached after the first read. The agent sends the full codebase once, the server holds onto that cached context, and every subsequent question only adds the new tokens on top.
In testing with a 16,000-token cached context, the first question took about 21 seconds on the Mac and 8.3 seconds on the Sparks. But the follow-up question, benefiting from the cache, dropped to about 3 seconds on the Mac and 1.4 seconds on the Sparks. The Sparks remain faster, but the gap shrinks and the expensive read only happens once per session rather than on every single message. The pain point reappears when the context keeps changing substantially: switching to a different repository, adding large new files, or running a session long enough to outgrow the cache.
Is the M5 Ultra or dual DGX Spark setup worth it for local LLMs?
It depends on the shape of your workload. If you’re mostly working with short prompts and care about fast, fluid output (chat-style use, iterative coding with cached context, lightweight agents), the M5 Ultra’s memory bandwidth and single-box simplicity make it a strong, often cheaper-to-operate choice, especially when paired with MLX instead of llama.cpp. If your work involves feeding in large codebases, long documents, or growing contexts regularly, and you need many concurrent users served efficiently, the dual Spark cluster’s compute advantage in prefill and its better scaling past four simultaneous users make it the stronger pick, despite the added complexity of clustering two units over an RDMA cable.
Price is close enough that it shouldn’t be the deciding factor on its own: the tested Mac configuration ran about $14,000, while two Sparks plus a high-speed interconnect cable landed closer to $10,000 to $11,000 all-in.
Frequently Asked Questions
What models were used in this benchmark?
The comparison used DeepSeek V4 Flash (roughly 150 to 160GB in size, requiring both Sparks to fit) and Qwen 3.8 Flash Next, both run in 4-bit quantized form on each platform.
Why does the DGX Spark need two units instead of one?
Everyone else built a construction worker.
We built the contractor.
One file at a time.
UI, API, database, deploy.
A single DGX Spark has 128GB of memory, which isn’t enough to load larger models like DeepSeek V4 Flash (around 150 to 160GB) plus the extra memory needed for context. Two Sparks combined provide 256GB, matching the Mac Studio’s total memory.
Does the Mac Studio support the same software as DGX Spark?
Partially. The Sparks primarily run vLLM, NVIDIA’s preferred engine for multi-GPU tensor-parallel inference, though llama.cpp also works there. The Mac was tested with both llama.cpp and MLX, Apple’s own framework, which produced notably faster generation speeds on some models.
Which setup handles multiple simultaneous users better?
The dual Spark cluster scales better under concurrency. In testing, the Mac’s total token throughput on DeepSeek peaked around four users (about 66 tokens per second) before dropping at eight users (46 tokens per second), while the Spark cluster kept increasing throughput as more users were added.
Is prompt processing speed something most users will notice?
Only with long inputs or cold sessions. At short prompt lengths (around 2,000 tokens), both machines finish in a second or two, a difference most users won’t notice. The gap becomes significant at tens of thousands of tokens, such as when loading an entire codebase, though prompt caching reduces this cost to a one-time hit per session rather than a per-question penalty.



