Self-hosting

DGX Station for Windows: Trillion Parameters on a Desk, Price TBD

·

NVIDIA's DGX Station for Windows hits the market this quarter, and it comes with the most aggressive memory spec ever put on a desk: up to 748GB of unified memory, built to run AI models with as many as one trillion parameters locally. NVIDIA announced it at GTC Taipei on June 1, 2026, and with ASUS, Dell, GIGABYTE, HP, MSI, and Supermicro all listed as system builders, the buying decisions are about to get real. Here is what the machine actually is, what "1 trillion parameters on a desk" really means, and who should care.

Short answer

  • DGX Station for Windows is a desktop AI supercomputer built on the Grace Blackwell Ultra desktop superchip: 72-core Grace CPU + Blackwell Ultra GPU, connected via NVLink-C2C. Announced June 1, 2026, shipping Q4 2026 from six OEMs. No price announced yet — that's not an oversight, it's the open question this page can't answer for you.
  • The spec that matters: 748GB unified memory and 20 PetaFLOPS FP4. NVIDIA's claim: it runs models up to 1 trillion parameters and hundreds of concurrent agents locally.
  • Our read: 748GB is enough for trillion-parameter Mixture-of-Experts models at low precision — not a 1T dense model (the weights alone would need roughly a terabyte at FP8; the math is below). This is a machine for MoE-era models, and that's exactly what the frontier looks like now.
  • Who it's for: enterprises standardized on Windows that want agents running next to their data; teams doing high-throughput inference or data science on datasets that don't fit a workstation. Who it's not for: individual developers — yet. Wait for the price, then run the rent-vs-buy math we did for the smaller DGX Spark.

The specs, from the horse's mouth

Everything below comes from NVIDIA's own announcement:

  • Compute: Grace Blackwell Ultra desktop superchip — 72-core Grace CPU paired with a Blackwell Ultra GPU over NVLink-C2C, rated at up to 20 PetaFLOPS FP4.
  • Memory: up to 748GB unified, addressable by both CPU and GPU — the number that breaks the usual desktop memory ceiling.
  • Networking: ConnectX-8 SuperNIC at up to 800Gb/s, so multiple Stations can be clustered for bigger workloads.
  • Pairing option: an additional RTX PRO Blackwell workstation GPU for visualization and simulation alongside the AI compute.
  • Software angle: NVIDIA OpenShell, an open-source agent runtime built on new Windows security and isolation primitives — each agent gets its own sandbox, and security policy is enforced at the system layer, not by prompting.
  • Availability: Q4 2026, from ASUS, Dell Technologies, GIGABYTE, HP, MSI, and Supermicro. Workloads can scale to Grace Blackwell Ultra in the data center or cloud.

What "1 trillion parameters" actually buys you

NVIDIA says the Station runs models with up to 1 trillion parameters locally. Take that claim seriously, but read it precisely. A 1T dense model at FP8 needs roughly a terabyte just for weights — more than the 748GB on offer. What fits is the current generation of trillion-class MoE models, where only a fraction of parameters activate per token, plus low-precision quantization. That is not a dodge: the models people actually want to run locally — the DeepSeek V4 family, Kimi K3, the Qwen MoE line — are exactly this shape. The Station is sized for the models that exist, not a hypothetical dense giant.

The other number worth staring at is 748GB of unified memory. For data science work, that means datasets that would normally require a server can sit in memory next to the GPU. For inference, it means the model, the KV cache, and your context live in one addressable pool — which is precisely the bottleneck that caps local deployment of large models on consumer hardware.

Who should actually consider one

Windows-standardized enterprises building agent fleets. This is NVIDIA's stated target and it's coherent: the Station runs hundreds of concurrent agents in isolated sandboxes, managed through the Windows compliance and device-management stack IT teams already use, with Linux workloads supported via WSL. For organizations that won't put Linux servers next to payroll data, this is the first machine that offers frontier-model compute inside their existing trust boundary.

Teams whose data can't move. High-throughput inference and data science on datasets that are too large, too sensitive, or too legally tangled to send to an API. The 748GB pool is the whole argument.

Nobody else yet. For individual developers and small teams, the honest answer is that no local purchase decision exists until there's a price. The smaller DGX Spark made sense or didn't depending on the rent-vs-buy math — we ran those numbers separately, and the Station will need the same treatment once OEM pricing lands. Cloud rental of GB300-class compute remains the flexible default, and for pure API workloads, renting by the token still wins on almost every horizon we've calculated.

The OpenShell detail that's easy to miss

The agent runtime might matter more than the silicon. OpenShell enforces security policy at the system layer — separate sandboxes per agent, infrastructure-level policy execution that the agent cannot override — instead of relying on system prompts to keep agents in line. If your compliance team has ever asked what happens when an agent with database access goes off-script, that architecture answer is the difference between a demo and a deployment. It's also the clearest sign of who NVIDIA is selling to here: not hobbyists, but IT departments.

The price is the story that hasn't happened yet

NVIDIA hasn't published a price, and neither have the six OEMs. Historical positioning says this sits well above consumer workstations, but that's inference, not reporting — the number doesn't exist publicly yet. When it does, the comparison that matters is against two alternatives: renting GB300-class capacity in the cloud, and the API route that costs nothing up front. Until then, the rational move for most teams is to keep watching this page's siblings: we've already run the rent-vs-buy math on the DGX Spark, and the self-hosting hardware trade-offs for running budget models like GLM-5.3-Flash on your own gear. The Station's math gets done the day a price exists.

Caveat: Specified October 8, 2026, against NVIDIA's official announcement post (GTC Taipei, June 1, 2026) — fetched and checked the same day. The 748GB / 20 PFLOPS FP4 / 72-core / 1T-parameter / 800Gb/s figures and the Q4 2026 window and OEM list all come from that announcement. No pricing has been announced by NVIDIA or any OEM as of this date. The FP8 weight-size arithmetic and the MoE-fit analysis are ours, derived from public model architecture disclosures, and labeled as such. Specs can shift between announcement and OEM launch — verify against the final vendor configuration sheets before buying.

Sources: NVIDIA Blog — DGX Station for Windows brings a trillion-parameter AI supercomputer to every enterprise desk (June 1, 2026)