Product Spotlight: The NVIDIA DGX Spark — a Personal AI Supercomputer for Developers
It's built around a genuinely different bet than a gaming GPU: trade raw memory bandwidth for a pool of unified memory big enough to hold models a consumer card can't touch at all.
The DGX Spark is a small desktop unit, close in footprint to a Mac mini-sized machine rather than a full tower, built around NVIDIA's GB10 Grace Blackwell Superchip — a single package combining a 20-core Arm-based Grace CPU with a Blackwell GPU carrying NVIDIA's fifth-generation Tensor Cores and fourth-generation RT cores. It launched October 15, 2025, after a delay from its originally planned May release.
The spec that defines the whole product is its memory: 128GB of LPDDR5x unified memory, shared directly between the CPU and GPU sides of the chip, at up to 273 GB/s of bandwidth, alongside up to 1 petaFLOP of AI performance at FP4 precision, 1TB or 4TB of self-encrypting NVMe storage, and ConnectX-7 networking supporting up to 200 Gbps.
That memory pool is what actually buys a developer something a consumer GPU can't: NVIDIA states a single DGX Spark can run AI models with up to roughly 200 billion parameters entirely on one unit, and its ConnectX-7 networking is specifically built to let two DGX Spark units link together to handle models up to about 405 billion parameters combined — a scale of model well out of reach of any consumer GPU's VRAM ceiling, including the 32GB top-end RTX 50 series card covered elsewhere in this category.
Why NVIDIA built this specific product, rather than just shipping a bigger consumer GPU, comes down to NVIDIA's own stated positioning: DGX Spark runs the same CUDA and Blackwell software stack as NVIDIA's full rack-scale DGX systems used in production data centers, so a developer can prototype and validate a workload locally on a desk, then move that exact same code to a much larger DGX deployment without a re-platforming step in between. It's a smaller, cheaper on-ramp to the same architecture, not a scaled-down, incompatible little cousin of it.
The honest trade-off deserves to be stated plainly rather than glossed over: DGX Spark's 273 GB/s of memory bandwidth is meaningfully lower than a dedicated gaming GPU's dedicated GDDR7 memory — the RTX 5090's roughly 1,792 GB/s, for comparison. That's the real shape of the bet DGX Spark makes: it trades raw memory throughput for enormous, unified capacity. A workload that's bottlenecked by whether a large model fits in memory at all benefits enormously from DGX Spark; a workload that already fits comfortably on a smaller card and just wants the fastest possible token-by-token throughput may actually be better served by that card instead.
On price, DGX Spark launched at $3,999 and has since risen to $4,699 — well above a high-end consumer GPU on cost alone, and positioned squarely as a purpose-built tool for developers doing serious local large-model work, prototyping against genuinely large models, or validating a workload before it moves to a full DGX-based production deployment, rather than as a general-purpose developer workstation or a gaming machine that happens to also run AI workloads. For a developer whose actual work is small-model prototyping or everyday development, the CPU and GPU hardware covered elsewhere in this category is the far more cost-effective starting point. DGX Spark earns its price specifically for the developer who has already hit a real capacity ceiling on that hardware and needs the unified memory pool to go meaningfully further.
Source: NVIDIA — DGX Spark Product Page
