A cheaper tier, a pricier flagship

The new box keeps everything that matters from the original: the GB10 Grace Blackwell Superchip, a 20-core Arm CPU, ConnectX-7 networking, 273 GB/s of unified memory bandwidth, DGX OS, and Nvidia's full AI software stack. The only thing that changed is capacity, halved to 64GB of unified memory shared between CPU and GPU.

Nvidia says one 64GB unit can run models up to 100 billion parameters locally, covering inference, fine-tuning, and agent workloads without cloud round-trips. The stack ships with the Agent Toolkit, CUDA-X libraries, Nemotron open models, and the usual local runtimes: Ollama, vLLM, llama.cpp, LM Studio, and PyTorch with CUDA.

The part of the announcement Nvidia talked about less is the price of the existing model. The 128GB DGX Spark debuted at $3,999, got bumped to $4,699 back in February, and now lists at $6,950. The stated cause is memory supply. Micron has warned DRAM stays tight through 2028, and the price ladder reads like a supply-chain chart, not a product refresh.

The cluster pitch, priced per unit

The genuinely new idea here isn't the smaller box. It's the NVIDIA Sync Cluster Assistant: a direct QSFP cable between two DGX Sparks' ConnectX-7 ports, no switch needed, pooling the two into 128GB of unified memory. Nvidia says the pair can then handle up to 200 billion parameter models, and that in its own test the clustered pair delivered up to 1.7x the performance of a single 128GB system on Qwen 3.8 27B.

Now run the numbers. Two 64GB units at $4,999 each is $9,998 for the same 128GB memory pool that one $6,950 box provides. You get double the compute and double the bandwidth for the $3,048 premium, but not cheaper memory. The "up to 1.7x" figure is one model, Nvidia's own benchmark, with no independent verification yet. Independent coverage has framed the 64GB unit as realistically suited to mid-size models in the 26 to 35 billion parameter range, a fair distance from the 100 billion headline.

So the honest reading: this is Nvidia managing scarcity, not generosity. A cheaper SKU softens the sticker shock while the flagship's real price quietly moves in the opposite direction.

Who it's actually for

Developers who want a local lab that doesn't phone home. Nvidia's own pitch lists round-the-clock coding or research agents, offloading inference from everyday laptops, and scaling to a second unit as workloads grow. A Sync Model Launcher due at month-end promises one-click launches of Qwen 3.8 27B on one box or a cluster, with OpenCode integration from a browser. Blender support is on the way for the creator side.

If you're pricing one of these, price per gigabyte of memory, not per box. The 64GB unit at $4,999 works out to roughly $78 per GB of unified memory. The 128GB unit at $6,950 works out to roughly $54 per GB. Cheaper entry price. More expensive memory. That's the whole story.

Sources

  1. [1] Tom's Hardware (via AI Weekly)Read source
  2. [2] StorageReviewRead source
  3. [3] NeoTeoRead source
  4. [4] ThePCEnthusiastRead source
  5. [5] TechnobezzRead source
  6. [6] ParticleRead source
  7. [7] GadgetBondRead source