Skip to content
CSuite
Sep 14, 202615 min read

The best PC for local AI on Windows in 2026: buy the right shape first

A tower stops at 32 GB, a mini PC reaches 96 GB, and the desktop with the AI sticker runs neither. Pick the shape before the SKU.

On this page
Model memory a Windows desktop can offer · orderable, September 2026
Three shapes of Windows PC. Pick by the memory rung, not the badge.
Memory
Runs at 4-bit
GPU tower
Unified mini PC
NPU “AI PC”
96 GB
120B-class MoE
·
·
64 GB
70B, comfortably
·
·
48 GB
70B, tight
·
·
32 GB
32B with room
·
·
24 GB
27–32B
used / last gen
·
·
16 GB
14B, long context
·
·
12 GB
12–14B
·
·
8 GB
8B chat models
·
·
~3 GB
built-in Windows models
Tower rungs are the memory sizes NVIDIA and AMD ship on current desktop cards; the dashed 24 GB rung exists only on last-generation or used cards. Mini PC rungs are what a 64 GB or 128 GB Strix Halo box can hand its GPU. The NPU column is the on-device models Windows itself ships, which is the only class an NPU is built for.

Walk into any electronics store in September 2026 and the desktop with the biggest “AI” sticker is the one least able to run a serious AI model. The sticker means the machine has a neural processing unit, a small chip built to run the little models Windows itself ships. It says nothing about the number that decides whether a 14B, 32B, or 70B model will load at all: how much memory the graphics processor can reach.

That number is why the Windows buying decision is a choice between three machine shapes, not a choice between brands. A tower with a discrete graphics card gives you a fixed ladder of memory sizes that tops out at 32 GB. A mini PC built on AMD’s Strix Halo chip shares one pool of up to 128 GB between its processor and graphics, and is the only sane Windows path past 32 GB. And a Copilot+ “AI PC” is a fine office computer that is the wrong tool for the models people mean when they say local AI. Pick the shape first. The exact SKU is a detail.

This guide is for people who want to buy a finished desktop and turn it on. If you would rather build one from parts, the build guide covers that path with the same memory logic. Every price below was read from a vendor page or a price tracker on September 13 or 14, 2026, and is linked where it appears. Vendors are compared on published facts only.

Three shapes of Windows PC, and the badge is on the wrong one

A language model has to sit entirely in memory the graphics processor can read, and it reads the whole thing for every word it generates. That single fact, explained in the RAM and VRAM post, sorts every Windows desktop into one of three shapes.

The GPU tower. A conventional case with a discrete NVIDIA or AMD card. The card has its own memory, fast and fixed, and the model has to fit in it. Your 64 GB of system RAM does not count. Towers are the fastest shape per token and the best value up to 16 GB, and they run every kind of model, including image and video generation, without fuss.

The unified-memory mini PC. A book-sized box built around AMD’s Ryzen AI Max+ 395, the chip everyone calls Strix Halo, where the processor and graphics share one pool of soldered memory. On a 128 GB box you can hand the graphics side up to 96 GB. Nothing in a consumer tower comes close, but the memory is slower, so big models run at a walking pace.

The NPU “AI PC.” A slim desktop or all-in-one with a processor whose neural unit clears Microsoft’s 40 trillion operations per second bar for the Copilot+ badge. The NPU accelerates Windows’ own built-in models and is a genuine efficiency win for those. It is not a place to run a 12B to 70B model, and the guide comes back to why below.

The three shapes, left to right: a tower with a discrete card, a unified-memory mini PC, and a slim office desktop of the kind that carries the Copilot+ badge. Only two of them are built for the models this post is about. Illustration generated with Seedream 5 Pro via Runware.

The one-line version: buy a tower if your ambition stops at 32B, buy a Strix Halo box if it does not, and buy the AI PC for reasons other than local AI.

A tower is a VRAM ladder with a missing step at 24 GB

NVIDIA’s current desktop lineup ships in exactly four memory sizes. On NVIDIA’s own comparison table, the RTX 5050, 5060, and one 5060 Ti carry 8 GB; the 5070 carries 12 GB; the 16 GB 5060 Ti, the 5070 Ti, and the 5080 carry 16 GB; and the 5090 carries 32 GB. There is no 24 GB card in the generation. AMD’s consumer side stops at 16 GB with the RX 9070 XT, and its 32 GB entry is a workstation card, the Radeon AI PRO R9700, which is the same silicon as the 9070 XT with double the memory at a $1,299 list price.

That missing 24 GB step matters because 24 GB is the rung where 27B to 32B models become comfortable. A 32B model at 4-bit is roughly 20 GB of weights plus context. On 16 GB it does not fit; on 32 GB it fits with room to spare. The gap between those two rungs on the price list is enormous, which is why the build guide still sends value buyers to a used 24 GB RTX 3090 or 4090. Prebuilt buyers do not get that option. A new tower either stops at 16 GB or jumps to 32 GB.

The desktop card ladder, sorted by memory then bandwidth
Memory decides what fits; bandwidth (the bar) decides how fast it generates. Power figures are the vendor’s card power and minimum system supply. Street prices are the lowest tracked US listing on September 13, 2026; the asterisk marks the 8 GB tracker price, since the 16 GB card was not separately tracked that day.
CardMemoryGB/sCard / PSUStreetMSRP
RTX 5060 Ti 16 GB16 GB
448
180 W / 600 W$680*$429
RTX 507012 GB
672
250 W / 650 W$887$549
RX 9070 XT16 GB
640
304 W / 750 W$799$599
RTX 5070 Ti16 GB
896
300 W / 750 W$1,200$749
RTX 508016 GB
960
360 W / 850 W$1,299$999
Radeon AI PRO R970032 GB
640
300 W / 750 W$1,299 list$1,299
RTX 509032 GB
1792
575 W / 1000 W$5,997$1,999

The street column is the part that has changed since July. NVIDIA’s list prices are still the launch numbers, but the memory shortage has pushed the RTX 5090’s lowest tracked US price to $5,997 on September 13, three times its $1,999 list. The 5080 sat at $1,299, the 5070 Ti at $1,200, the 5070 at $887, and even the 8 GB 5060 Ti at $680, 79% over list. AMD’s 9070 XT was $799 against a $599 list.

The number printed on the box is not the number that matters. What matters is the memory soldered around the chip, and that is fixed for the life of the card. Photo by Christian Wiediger on Unsplash.

Two consequences for a buyer. First, the 16 GB tier is where the value lives: a 5070 Ti or a 9070 XT holds a 14B model with a long context window and runs image models comfortably, at a quarter of the 5090’s street price. Second, at the 32 GB tier the card alone now costs more than a complete prebuilt with the same card, which flips the usual advice. If you want a 5090, buy it inside a tower.

The prebuilt lines, by what the slot holds

The big three Windows vendors each sell a gaming tower that doubles as a local AI box, and the useful way to read their catalogs is by the top card each chassis accepts.

Dell. Dell’s Alienware page lists the Aurora from $1,849.99 with cards up to an RTX 5080, which makes it a 16 GB machine however you configure it. The larger Area-51 starts at $4,199.99 on Intel and $4,099.99 on AMD and goes up to the RTX 5090. That split is the whole story of the lineup: the mid-tower is a 16 GB box, the full tower is the 32 GB box.

HP. HP’s OMEN desktop page draws the same line. The OMEN 35L tops out at an RTX 5080, the small OMEN 16L at an 8 GB 5060 Ti, and only the OMEN MAX 45L reaches the 5090, with an RX 9070 XT as the AMD option. A May 2026 review put the MAX 45L’s starting price at $3,199, and a 5090 build with 64 GB and 4 TB at $6,500 outside of sales.

Lenovo. The Legion Tower 7i Gen 10 is Lenovo’s 5090 chassis; the configuration on Amazon’s listing pairs the card with 64 GB of RAM and a 1200 W supply, which is the right amount of both. Lenovo’s smaller Tower 5i lines stop at the 5070 Ti.

Prices on all three move weekly and the sale price is usually the real price. What does not move is the ceiling per chassis, so decide 16 GB or 32 GB first and only then open the configurator. Expect base prices to drift up: in January IDC relayed that Lenovo, Dell, HP, Acer, and Asus had warned of 15 to 20 percent increases as AI data centers absorbed the memory supply, and TrendForce still expects DRAM contract prices to climb another 13 to 18 percent this quarter, after a 59.5 percent jump last quarter.

Past 24 GB, the Windows path is a mini PC with unified memory

AMD’s Ryzen AI Max+ 395 is a laptop-class chip that turned out to be the most interesting desktop part of the year. Per AMD’s spec page, it pairs 16 cores with a 40-unit Radeon 8060S graphics engine and up to 128 GB of LPDDR5x on a 256-bit bus. The graphics side does not have memory of its own; it borrows from that pool. AMD’s Variable Graphics Memory FAQ puts the ceiling at 96 GB on a 128 GB system and the bandwidth at 256 GB/s, and you set the split in the Adrenalin driver software on Windows.

Read those two numbers against the tower ladder. 96 GB is three times what an RTX 5090 holds, so a 70B model at 4-bit (about 43 GB) fits with room for context, and a 120B mixture-of-experts model (about 65 GB) fits too. But 256 GB/s is one seventh of the 5090’s 1,792 GB/s, and generation speed tracks bandwidth. The box fits models the tower cannot, and runs the models both can fit at a fraction of the speed.

Published generation speed, tokens per second
Violet is a 128 GB Strix Halo mini PC (HP Z2 Mini G1a, Ollama and LM Studio, Gemma 3 and gpt-oss). Amber is a desktop RTX 5090 (llama.cpp, Qwen3). Different models and harnesses, so read them as ballparks. Anything above 20 is faster than most people read.
7–8B
fits both
37
186
14B
fits both
19
124
32B dense
fits both, 5090 at 32 GB
9
61
70B dense
does not fit 32 GB
4
n/a
120B MoE
does not fit 32 GB
40
n/a

The figure uses StorageReview’s HP Z2 Mini G1a numbers for the mini PC: 36.86 tokens per second on a 7B model, 9.38 on a 32B, 4.24 on a 70B, and close to 40 on gpt-oss-120b, whose mixture-of-experts design only touches a few billion parameters per token. For the tower it uses Hardware Corner’s RTX 5090 runs: 185.91 on Qwen3 8B, 123.79 on 14B, 61.38 on the dense 32B. The 70B and 120B rows are blank for the tower because they do not fit in 32 GB, and that blank is the entire argument for the shape.

A Strix Halo mini PC is the size of a hardcover book and holds more model memory than any graphics card you can put in a tower. Illustration generated with Seedream 5 Pro via Runware.

Three boxes carry the chip in a desktop shell. GMKtec’s EVO-X2 was listed at $2,199.99 on sale from $2,599.99 on September 14, shipping with Windows 11 Pro. HP’s Z2 Mini G1a is the workstation take, with ECC memory and the same 96 GB graphics ceiling; StorageReview’s tested 128 GB unit was priced at $3,342.65. And Framework’s Desktop tells the memory-shortage story in one line: the 128 GB model went from $1,999 to $2,459 in January, and Framework’s page listed it at $3,449 and out of stock on September 14. Buy 128 GB, not 64 GB: the 64 GB box hands its GPU about 48 GB, which holds a 70B model with no room to breathe.

DGX Spark is a developer box, not a faster mini PC

NVIDIA sells its own book-sized 128 GB machine, and it is tempting to read it as the Strix Halo box with CUDA. The DGX Spark spec page lists a Grace Blackwell GB10 chip, 128 GB of unified memory at 273 GB/s, 4 TB of storage, up to 1 petaflop of 4-bit compute, and inference on models up to 200 billion parameters, with four units linkable for 700 billion. The bandwidth is in the same class as Strix Halo, so token speed on a big dense model will be too.

Two things keep it out of a Windows buying guide. It does not run Windows: the operating system is NVIDIA’s DGX OS, an Ubuntu derivative, and the processor is Arm, so your Windows software does not come along. And the price moved the wrong way. NVIDIA raised the Founders Edition from $3,999 to $4,699 in February, a change staff confirmed on NVIDIA’s own forum and Tom’s Hardware tied to the memory shortage. It is the right box for someone fine-tuning models in the CUDA toolchain who wants the same stack as the data center. It is the wrong box for someone who wants a Windows desktop that runs a 70B model.

The other pro path is a workstation card in a tower. NVIDIA’s RTX PRO 4500 Blackwell gives you 32 GB of error-corrected GDDR7 at 896 GB/s in a 200 W card, and the 96 GB RTX PRO 6000 exists at a list price Tom’s Hardware reported NVIDIA doubled to $16,000. For almost everyone, a 128 GB Strix Halo box at one-fifth that price is the sensible way to reach 96 GB.

A Copilot+ NPU serves the built-in models, not the 12B to 70B class

Microsoft’s Copilot+ PC developer guide defines the class as Windows 11 hardware with an NPU that can perform more than 40 trillion operations per second, built for the features Windows ships on it: real-time translation, image generation, and the on-device language model behind Windows’ own text tools. The NPU runs those at low power, which is the whole point on a laptop. On a desktop the battery argument disappears, but the features work.

What the NPU is not built for is loading a multi-gigabyte model of your choosing. Microsoft’s own on-device model, Phi Silica, is the clearest evidence. Its documentation says it runs on the NPU on Copilot+ PCs and, since a June 2026 preview, also on ordinary Windows 11 machines with a discrete GPU: a GeForce RTX 30 series or newer, or a Radeon RX 9060 or newer, with 6 GB of memory. Read that the other way round. A 6 GB gaming card from 2020 qualifies to run Microsoft’s flagship on-device model. The NPU is a lower-power way to run the same small model, not a bigger one. None of the runtimes in the next section, Ollama, llama.cpp, or the app this site belongs to, use the NPU at all; they run on the GPU.

Copilot+ desktops do now exist. AMD brought its Ryzen AI PRO 400 series to the AM5 desktop socket in March 2026 with a 50 TOPS NPU that clears the badge threshold. The useful reading of that launch is the socket, not the NPU: an AM5 board with a 50 TOPS chip is a perfectly good tower, and the moment you drop a 16 GB card into it, the card does the local AI work and the NPU keeps doing the Windows features. Buy the badge if you want the Windows features. Buy the graphics card for everything else.

The Windows layer: 25H2, one driver, and the right backend

The Windows version. Microsoft’s release information page shows three versions in service this month, and one of them is about to fall off. Windows 11 24H2 Home and Pro editions reach end of updates on October 13, 2026. 25H2 is the version an existing machine should be on, and a new desktop should ship with it. 26H1 is a special case: Microsoft scopes it to new devices coming to market in early 2026 and does not offer it as an in-place update to a 24H2 or 25H2 machine.

The driver. NVIDIA ships two driver tracks for the same cards, and its own drivers page describes the split plainly: Game Ready for people who want day-one support for new games, Studio for people who prioritize reliability in creative applications. Either track runs every runtime below; Studio is the calmer choice for a machine that runs models for hours. Ollama’s hardware page asks for driver 550 or newer on NVIDIA. On AMD, install Adrenalin directly from AMD rather than trusting the version Windows Update offers; Microsoft’s own Phi Silica docs make the same point about OEM drivers being too old.

The backend. A model runtime talks to the graphics card through one of a few interfaces, and on Windows the choice is made for you by the card. CUDA is NVIDIA’s own and the best-supported path. Vulkan is the cross-vendor graphics interface that llama.cpp and Ollama can also use for compute; it runs on NVIDIA, AMD, and Intel cards alike. ROCm, AMD’s answer to CUDA, is more selective on Windows than people expect: Ollama’s hardware page lists the Radeon RX 7000 and PRO W7000 families for Windows ROCm, while the RX 9000 cards and the R9700 appear only on its Linux list. On Windows those cards run through Vulkan, which Ollama enables by default when the backend is installed. That is not the handicap it sounds like. One March 2026 test of a 9070 XT had llama.cpp on Vulkan at 62 tokens per second on a 9B model, ahead of a ROCm stack that lacked native kernels for the card. DirectML, the older Windows-native path, is being replaced by Windows ML per Microsoft’s Copilot+ guide and is not what these runtimes use.

Which backend actually runs your model on Windows
From Ollama’s hardware page and llama.cpp’s build docs, September 2026. CSuite bundles Ollama and ships CUDA only when it finds an NVIDIA GPU; the Vulkan backend is always included. Nothing in this row set touches the NPU.
GPUOllamallama.cppCSuite
NVIDIA GeForce / RTX PROCUDA (driver 550+)CUDA or VulkanCUDA
AMD Radeon RX 7000ROCm (on the Windows list)HIP or VulkanVulkan
AMD Radeon RX 9000 / R9700Vulkan (ROCm is Linux-only here)Vulkan, or HIP if you build itVulkan
Strix Halo Radeon 8060SVulkanVulkanVulkan
Copilot+ NPUnot usednot usednot used

CSuite, the desktop app this blog belongs to, bundles Ollama and makes the same choice automatically: on Windows it keeps the CUDA runtimes only when it finds an NVIDIA GPU, and it always ships the Vulkan backend, so an AMD or Intel card is accelerated without a setting. The Windows build installs the same way on all three shapes; on the NPU desktop it simply runs on the integrated graphics, slowly. The Windows apps roundup covers the alternatives.

WSL2 for Linux-only tooling. Some training and serving tools ship for Linux first. Windows Subsystem for Linux passes the GPU through, and NVIDIA’s CUDA on WSL guide is emphatic about the one rule: install the Windows NVIDIA driver and nothing else, and “do not install any Linux display driver in WSL.” Microsoft’s WSL GPU page adds that a machine with several GPUs can expose only one to WSL at a time. For day-to-day model running you do not need WSL; Ollama and llama.cpp are native Windows programs.

The rest of the box: RAM, power, and why two GPUs rarely pencil out

System RAM. On a tower the model lives in the card’s memory, so system RAM is not the headline number. It still matters: a model that slightly overflows VRAM can spill layers into RAM and keep running, slowly. 32 GB is the floor for a 16 GB card and 64 GB is the right pairing for a 32 GB card, which is what the Lenovo and HP 5090 configurations ship. Do not pay for 128 GB of system RAM to “help” a graphics card; it does not.

A 430 W supply and a short card: fine for an office PC, and the reason many older towers cannot take a serious upgrade. A single RTX 5090 asks for a 1000 W system. Photo by Luke Hodde on Unsplash.

Power and airflow. NVIDIA’s spec pages are the source of truth here: the RTX 5090 draws 575 W on its own and asks for a 1000 W system supply; the 5080 wants 850 W and the 5070 Ti 750 W. A prebuilt has been specified for the card it ships with, so the question only bites when you plan to upgrade later. A mid-tower sold with a 5070 Ti and a 750 W supply will not take a 5090 without a new supply and, often, a bigger case. Generation runs the card at full load for minutes at a time, longer than most games, so a case with a front intake and a top exhaust is worth more than lighting.

Two cards. The idea of pairing two 16 GB cards to make 32 GB comes up constantly and rarely survives the arithmetic. The runtimes can split a model across two cards, but the second card adds its power draw, its slot, and its price, and the two halves talk over the motherboard rather than a single memory bus. Two 5070 Ti at street price cost nearly twice a 32 GB R9700 at list, draw twice the power, and need a board with two properly wired slots that almost no prebuilt has. If you need 32 GB, buy it on one card, or buy the mini PC.

The shortlist at three budgets

These are shapes and card tiers rather than exact SKUs, because the SKU in stock next week will be different and the shape will not. Check your own model ambitions against the calculator before you spend.

Around $2,000 · GPU tower
Prebuilt with an RTX 5070 Ti or RX 9070 XT
16 GB VRAM · 32 GB RAM
8B to 14B at speed; 20B MoE; image models
Every mainstream line stops at 16 GB in this band, so buy the fastest 16 GB card and skip the 12 GB 5070.
Around $3,000 · Unified mini PC
Strix Halo box at 128 GB (EVO-X2, Z2 Mini G1a, Framework)
96 GB to the GPU · 256 GB/s
70B dense slowly; 120B MoE at reading speed
The only Windows shape past 32 GB without a five-figure card. Slower per token than a tower, but it fits.
$5,000 and up · GPU tower
Prebuilt with an RTX 5090 (Area-51, OMEN MAX 45L, Legion Tower 7i)
32 GB VRAM · 64 GB RAM · 1000 W
32B dense at 60+ tokens/s; everything smaller, fast
The fastest consumer path to 32 GB. The card alone is near $6,000 on the street, so the prebuilt is the better buy.

Two things stayed constant across every page read for this post. The 24 GB rung, the sweet spot for 27B to 32B models, does not exist on a new Windows desktop card, so a buyer chooses between 16 GB and 32 GB and the chassis decides which. And memory is the part getting more expensive, not less, so the box you buy now is the box you will run for years. Decide the model size you want to run, find its rung in the hero, and buy the shape that owns that rung. The badge on the front of the case can say whatever it likes.

If you also carry a laptop, or your desk is a Mac, the same logic runs through the laptop guide and the Mac guide, where unified memory is the default rather than the exception.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app