NVIDIA calls DGX Spark a personal AI supercomputer. NetworkChuck put that claim under a more useful question: can the little GB10 box beat his dual-RTX 4090 AI server? His answer is no for the speed tests he ran. Yet the Spark can hold larger and more concurrent workloads in a way his two gaming GPUs cannot easily match. The decision is about which constraint your work hits first: memory or throughput.
- Text and images: Chuck measured 36 versus 132 tokens/second for a small model, and about one versus 11 iterations/second in his ComfyUI test. His dual-4090 server was faster in both.
- Memory: Spark has 128GB of coherent unified system memory. Chuck's server has two separate 24GB VRAM pools.
- What it means: Spark can fit larger local development jobs, but may finish smaller jobs more slowly. These are creator-observed results for particular setups, not universal benchmarks.
Watch NetworkChuck's DGX Spark Test
Credit and disclosure: NetworkChuck's original review, published 14 October 2025, and its supplied transcript. NVIDIA sent him the device; he says NVIDIA did not review or control the video. The video also includes a sponsored Twingate segment. This article is not sponsored. Correction: the chip is the GB10 Grace Blackwell Superchip, not “GP10.”
What DGX Spark Is Built For
NVIDIA's hardware guide lists a 20-core Arm CPU, Blackwell GPU, 128GB LPDDR5x coherent unified memory, 273GB/s memory bandwidth, 4TB storage on the Founders Edition, 10GbE, and a ConnectX-7 200Gb/s link for pairing systems. The 240W figure is the supplied power adapter rating, not a measured constant wall draw. “Up to 1 PFLOP” is a theoretical FP4 figure with sparsity, not the speed of an ordinary chat response.
NVIDIA says a single Spark can run inference on models up to 200B parameters and fine-tune models up to 70B. Those are capability ceilings under suitable model formats and workflows. Quantization, context/KV cache, runtime overhead, and software support determine what really fits; loading a model does not tell you whether its token speed is useful. NVIDIA opened orders on 15 October 2025, the day after Chuck's video.
Spark Versus the Dual-4090 Server
- Small-model text generation: Chuck reports 132 tokens/second on “Terry,” his two-4090 server, and 36 tokens/second on “Larry,” the Spark. His on-screen test is a snapshot of his configuration, not a standardized, repeatable benchmark suite.
- Llama 3.3 70B: Terry also wins the larger-model demonstration, but the video does not provide a clean, matched token-rate table or enough configuration detail to generalize the margin.
- ComfyUI image generation: Chuck reports around 11 iterations/second on Terry and roughly one on Spark. He notes that the test appeared to use one of Terry's GPUs. This pipeline favors the workstation for the measured job, even with that caveat.
- Small-model training: Terry takes about one second per iteration versus about three seconds on Spark in the training demo. Chuck begins loading a larger 70B workflow, but does not show a completed 70B fine-tune or its quality. Treat that as a capacity exploration, not a verified training result.
The lesson is not that the Spark “fails.” A dedicated RTX GPU has much faster local VRAM for jobs that fit. The Spark trades some of that throughput for a much larger shared memory space in a small package. Compare with the later Qwen3.8-27B test, where precision and memory path also changed between machines.
Why 128GB Changes the Workload
Chuck runs a multi-model setup that he reports using about 89GB of Spark memory. His two RTX 4090s each have 24GB of dedicated VRAM; their combined 48GB is not a transparent 48GB allocation for every program. Spark's 128GB is likewise not 128GB of dedicated GPU VRAM: CPU, operating system, containers, models, and context cache share it. Still, a single coherent pool makes certain large-model and multi-model experiments much simpler.
That is the useful distinction for a team that must run retrieval, embeddings, a larger reasoning model, and a draft model locally. It does not automatically make the workflow cheaper or private. Check data retention, model licenses, user access, patching, backups, and the quality of human-reviewed outputs before putting client information on any local machine.
FP4, Software, and Secure Access
The GB10's Blackwell Tensor Cores support FP4, and NVIDIA publishes Spark tutorials for local AI workflows. Quantizing a model can reduce memory use, but accuracy depends on the model, quantization method, and task. Chuck also demonstrates speculative decoding: a smaller draft model proposes tokens that a larger model verifies. That can improve particular inference runs, but his visual demo supplies no controlled token-rate measurement, so it is not proof that Spark reverses the earlier speed gap.
DGX OS comes preconfigured for NVIDIA's stack. NVIDIA Sync helps connect from a laptop and launch tools such as VS Code or Cursor through managed SSH and tunnels. For off-network access, use a controlled private-access method rather than exposing development ports to the public internet.
At 10:42, Chuck introduces Twingate through his sponsor link for remote access.
Twingate's pricing page currently lists a free Starter tier for up to five users; check the plan and access policy before depending on it. NVIDIA Sync also documents a Tailscale connection option. Neither tool removes the need for account controls and backups.
Who Should Buy One?
Consider Spark if you repeatedly need large local memory, develop or fine-tune models that will not fit on your current GPU, prefer NVIDIA's supported software path, and have measured enough work to justify a dedicated box. Choose a fast discrete GPU or cloud time if your main job is interactive chat, image generation, or occasional experiments on models that already fit. Rent first if you cannot name the workload and weekly utilization.
Chuck's 4TB Founders Edition was $3,999 at filming; that is historical launch context, not today's price. His annual electricity comparison assumes 24/7 operation and his own wattage and tariff. A 240W power-supply rating is not a power bill. For a purchase decision, benchmark your model, quantization, context length, concurrent users, wall power, and time to an accepted result, then compare full hardware and cloud costs. Our cloud-versus-home hardware guide offers a broader framework.
Turn This Into a Business Idea
The commercial opening is not “sell a Spark.” It is a repeatable workflow for a buyer whose data, model size, or review process makes local or hybrid AI useful. The prompt below asks for the buyer and evidence before any hardware commitment.
Find a buyer before buying the box
Test a local-first offer against real cost, data needs, and a manual baseline.
Video Chapters
| Time | Topic |
|---|---|
| 00:00 | Introducing the palm-sized Spark |
| 00:34 | GB10, 128GB memory, and specifications |
| 01:27 | Text inference versus dual RTX 4090 |
| 02:33 | Why Spark's strengths differ |
| 04:49 | ComfyUI image generation |
| 07:39 | Training and memory capacity |
| 09:26 | DGX OS and NVIDIA Sync |
| 10:42 | Sponsored Twingate remote-access segment |
| 12:45 | FP4 and speculative decoding |
| 14:47 | Price, power, and buyer fit |
| 18:33 | Chuck's purchase verdict |
Sources and Useful Links
- Creator test: NetworkChuck's original video. The measured speeds and opinions above are his, not official NVIDIA benchmarks.
- Hardware and availability: NVIDIA product page, hardware guide, and October 2025 launch announcement.
- Setup and access: NVIDIA Sync guide and Twingate plans. The source video's Twingate link is a sponsor link.
- More from NetworkChuck: his five-Mac-Studio AI computer and local voice-cloning project. These are related videos, not tests reproduced in this article.