NVIDIA launches 64GB DGX Spark with two-system clustering for larger local AI models and agents

NVIDIA has introduced a new 64GB DGX Spark configuration designed for developers running AI models and autonomous agents locally, extending the DGX Spark platform with a lower-memory entry point and a new path for scaling workloads across multiple systems.
The machine retains the GB10 Grace Blackwell Superchip, DGX OS, ConnectX-7 networking and NVIDIA's AI software stack. NVIDIA says a single 64GB DGX Spark can run models containing up to 100 billion parameters, depending on model architecture, precision, quantization and runtime requirements.
Two systems can also be connected directly and configured through NVIDIA Sync Cluster Assistant, providing 128GB of combined unified memory and supporting models with up to 200 billion parameters.
The 64GB configuration is scheduled to become available on October 23, 2026, through Acer, ASUS, Dell, Gigabyte, HP and MSI, with systems starting at $4,999.
··········
NVIDIA DGX SPARK 64GB AT A GLANCE
........
Specification | DGX Spark 64GB |
Unified memory | 64GB |
Processor | GB10 Grace Blackwell Superchip |
Operating environment | DGX OS |
Maximum stated model size | Up to 100B parameters |
Two-system memory | 128GB combined |
Maximum stated model size with two systems | Up to 200B parameters |
Networking | ConnectX-7 |
Two-node connection | QSFP |
AI frameworks | PyTorch, vLLM, Ollama and others |
Starting price | $4,999 |
Availability | October 23, 2026 |
........
The new configuration sits below the existing higher-memory DGX Spark while preserving the same Grace Blackwell platform. Developers can therefore begin with one 64GB machine and add a second node when larger models, longer contexts or additional concurrent workloads require more memory.
··········
THE 64GB SYSTEM RETAINS THE GB10 GRACE BLACKWELL ARCHITECTURE
The reduction to 64GB does not introduce a different processor platform.
DGX Spark continues to use the GB10 Grace Blackwell Superchip, combining Arm-based CPU resources and Blackwell GPU compute in a compact system built around unified memory.
Unified memory allows CPU and GPU resources to operate against a common memory pool rather than dividing workloads between conventional system RAM and discrete GPU VRAM. For local AI inference, this increases the amount of memory that can be allocated to model weights, context and runtime state within a single system.
The machine also retains DGX OS, CUDA-accelerated libraries and ConnectX-7 networking. Software developed on one DGX Spark can therefore remain within the same NVIDIA environment when a workload is expanded to a second machine.
··········
A SINGLE SYSTEM CAN SUPPORT MODELS UP TO 100 BILLION PARAMETERS
NVIDIA specifies support for AI models containing up to 100 billion parameters on one 64GB DGX Spark.
The parameter count should not be interpreted as a fixed memory equivalence. Model memory consumption changes substantially with numerical precision and quantization.
A 100B model represented at 16 bits per parameter would require roughly 200GB for weights alone and could not fit inside 64GB. At 4-bit quantization, the theoretical raw weight requirement falls to approximately 50GB before accounting for quantization metadata, runtime overhead, KV cache and other memory requirements.
Context length also affects practical capacity. Longer prompts and persistent agent sessions consume additional memory, while concurrent inference requests can multiply KV-cache requirements.
The 100B specification therefore describes the upper range achievable with appropriate model configurations rather than unrestricted execution of every model below that parameter count.
··········
TWO DGX SPARK SYSTEMS CREATE A 128GB LOCAL AI CLUSTER
Two 64GB DGX Spark machines can be connected through their ConnectX-7 interfaces, producing a two-node environment with 128GB of combined unified memory.
NVIDIA states that this configuration can support models containing up to 200 billion parameters.
The second machine can also be used to increase capacity for workloads that do not require a larger base model. A developer could allocate the additional resources to longer contexts, concurrent inference sessions or multiple autonomous agents rather than loading a single 200B model.
This creates a modular upgrade path: the original machine remains part of the expanded environment instead of being replaced when local workloads exceed 64GB.
··········
SYNC CLUSTER ASSISTANT AUTOMATES TWO-NODE CONFIGURATION
Distributed inference requires more than connecting two computers physically. Nodes need compatible network settings, connectivity validation and a runtime capable of distributing workloads between systems.
NVIDIA Sync Cluster Assistant
handles much of the initial configuration for DGX Spark clusters.
The software detects connected systems, verifies their configuration and prepares the ConnectX-7 network for communication between nodes. For a two-system installation, DGX Spark machines can be connected directly through QSFP rather than requiring a separate network switch.
ConnectX-7 provides the high-bandwidth link needed to exchange model data and intermediate results between the machines. Network communication still introduces overhead, meaning two systems do not automatically deliver twice the performance of one.
··········
NVIDIA REPORTS UP TO 1.7X PERFORMANCE WITH TWO SYSTEMS
NVIDIA tested a two-system configuration using Qwen 3.8 27B and reported performance of up to 1.7× that of a single DGX Spark.
........
Configuration | Unified memory | Maximum stated model capacity | Qwen 3.8 27B result |
1× DGX Spark 64GB | 64GB | Up to 100B | 1.0× baseline |
2× DGX Spark 64GB | 128GB | Up to 200B | Up to 1.7× |
........
Data Studios calculation: a 1.7× result corresponds to approximately 85% of ideal two-node linear scaling, where perfect scaling would produce 2.0× performance.
The result applies to NVIDIA's specific Qwen 3.8 27B test. Actual scaling depends on the model, inference framework, quantization, batch size and communication required between nodes.
For workloads exceeding the memory of one machine, the primary benefit can be increased capacity even when performance scaling remains below 2×.
··········
LOCAL AI AGENTS CAN USE THE EXTRA MEMORY FOR CONTEXT AND CONCURRENCY
DGX Spark is increasingly being positioned for persistent local AI agents in addition to conventional model inference.
Agentic workloads can require substantial memory even when the underlying model fits comfortably within the system. Long-running sessions accumulate context, parallel agents create separate inference states and simultaneous requests increase KV-cache consumption.
A 128GB two-node configuration can therefore be useful without approaching the 200B model limit.
Developers can use the additional capacity for several smaller agents, longer context windows or concurrent inference workloads while keeping processing on local hardware.
Local execution can also keep code repositories, documents and other working data within infrastructure controlled by the user or organization. It does not remove the need for permission controls, credential protection and isolation when agents are allowed to execute actions.
··········
DGX SPARK INCLUDES A COMPLETE LOCAL AI SOFTWARE ENVIRONMENT
The hardware ships within NVIDIA's existing AI development ecosystem.
DGX Spark supports CUDA-X libraries, NVIDIA Agent Toolkit, Nemotron models, PyTorch, Ollama and vLLM, providing several options for local inference, model experimentation and agent development.
NVIDIA is also preparing Sync Model Launcher, designed to simplify downloading and launching supported models across either one DGX Spark or a clustered configuration.
This reduces the amount of manual distributed-inference configuration required when moving a workload from one node to two.
The common software environment is particularly relevant for development workflows in which a model begins on a single machine and later requires additional local compute or memory.
··········
THE $4,999 CONFIGURATION LOWERS THE ENTRY POINT FOR DGX SPARK
The 64GB DGX Spark starts at $4,999, placing it firmly in professional development and workstation territory while reducing the initial hardware commitment required to enter the DGX Spark ecosystem.
The economics differ substantially from cloud inference.
Local hardware requires upfront capital expenditure, electricity and administration but does not generate a new per-token inference charge each time a locally hosted model runs. Cloud infrastructure requires less dedicated hardware while converting compute consumption into recurring operating expenditure.
For persistent agents or frequently used internal models, utilization becomes central to the comparison. A machine running occasional experiments has a different cost profile from hardware executing models continuously throughout the working day.
The ability to add a second DGX Spark also allows organizations to increase local capacity without discarding the first system.
··········
NVIDIA IS MAKING LOCAL AI COMPUTE MODULAR
The 64GB DGX Spark combines a smaller initial memory configuration with a scale-out architecture.
A single machine provides 64GB of unified memory and support for models up to NVIDIA's stated 100B-parameter ceiling. Connecting two machines increases combined memory to 128GB and raises the stated model ceiling to 200B, while NVIDIA's Qwen 3.8 27B test demonstrates that the second node can also increase inference performance.
The architecture can therefore scale along several dimensions: model size, context length, concurrent inference and the number of local agents.
Instead of treating a compact AI workstation as a fixed hardware configuration, NVIDIA is extending DGX Spark toward a modular local compute environment in which additional nodes can be introduced as workloads grow.
··········
FOLLOW US FOR MORE.
DATA STUDIOS
datastudios.org




