Skip to content
CSuite
ListicleLocal AIHardwareSep 8, 202614 min read

The best laptop for local AI in 2026: Mac, Windows, and Linux compared

A laptop graphics card stops at 24 GB and unified memory keeps climbing. One number per platform decides what your laptop will ever run.

Two memory ladders · laptops orderable, September 2026
A laptop graphics card stops at 24 GB. Unified memory keeps climbing.
Discrete VRAM
Model at 4-bit
Unified memory
8 GB
8B chat model
~5 GB
16 GB
12 GB
12–14B
~8–9 GB
24 GB
24 GB
27–32B dense
~17–20 GB
32–48 GB
no laptop card holds it
70B
~43 GB
64 GB
no laptop card holds it
120B-class MoE
~65 GB
96–128 GB
Unified tiers assume the GPU can use roughly three quarters of system memory, which is the default on Apple silicon and the maximum AMD lets you assign on Strix Halo. VRAM tiers are the RTX 50 laptop lineup as NVIDIA ships it. Sizes are 4-bit weights plus working context.

Walk into a laptop aisle in September 2026 and ask for “something that can run AI.” You will be handed a gaming laptop with an RTX badge, a fan the size of a saucer, and 8 GB of graphics memory. It will run an 8B model beautifully and a 32B model not at all. A silent MacBook Air with 32 GB, at a similar price, runs the 32B model. Neither machine is wrong. They are on different ladders, and the salesperson only knows one of them.

That is the whole problem with buying a portable machine for local AI this year. The decision is a single number, but the number has a different name on each platform: unified memory on Apple silicon and AMD’s Strix Halo chips, VRAM on any laptop with a separate graphics card. The two do not convert one to one. And because the AI boom has made memory the scarcest part in the whole box, the number you buy this year is the ceiling you live with for the life of the machine.

This guide is the cross-platform sibling of our Mac buying guide. It lays the two ladders side by side, shows what each rung runs, and ends with one pick per platform at three budgets. The short version: below about 16 GB, a discrete GPU is faster per dollar. Above it, a unified-memory laptop wins outright, because no laptop graphics card goes past 24 GB and unified memory keeps going to 128.

Memory is a cliff, and a laptop’s memory is fixed for life

A language model has to sit entirely in memory the GPU can reach while it runs. A model that needs 20 GB does not run a bit slower on a 16 GB machine. It either refuses to load or spills onto the SSD and produces a word every few seconds. That is why every table in this guide is sorted by memory rather than by benchmark score. The arithmetic behind it is one line, worked through in our memory-math explainer: a 4-bit model needs about 0.6 GB per billion parameters, plus a few gigabytes for the conversation itself.

On a laptop the cliff is permanent. Apple solders memory to the chip package and states plainly that RAM cannot be upgraded later. Strix Halo laptops do the same: ASUS lists the Flow Z13’s 32, 64, and 128 GB tiers as LPDDR5X 8000 onboard. And a graphics card’s VRAM is fixed by definition, even on a Framework Laptop 16 where the whole GPU module swaps out. The upgrade you skip at checkout is skipped for good.

The timing makes it worse. Memory makers spent 2025 converting consumer lines to data-center parts, and laptop makers were warned in December to expect list prices up at least 20 percent in 2026, with no relief before 2027. Apple made it concrete on June 25: the 13-inch MacBook Air went from $1,099 to $1,299 and the 14-inch M5 Max from $3,599 to $4,099, per EveryMac’s price-hike ledger. Framework charges $699 for an 8 GB RTX 5070 module and $1,199 for the 12 GB version of the same chip. Four gigabytes of memory now costs $500. Buy the memory deliberately.

Left, one pool of memory around one package: the GPU can use most of it. Right, a graphics module with its own small memory: the model has to fit in that cluster and nowhere else. Illustration generated with Seedream 5 Pro via Runware.

Two memory ladders, and they do not convert

Here is the conversion the aisle never explains. On a Mac, the CPU and GPU share one pool of memory on the same package, and macOS lets the graphics layer use roughly 75 percent of it by default. A 32 GB MacBook is, for this purpose, a 24 GB graphics card. A 64 GB MacBook Pro is a 48 GB one, and a 128 GB M5 Max is a 96 GB one, which is more than any consumer graphics card, desktop or laptop, has ever shipped with.

AMD’s Strix Halo chips, sold as Ryzen AI Max, work the same way on Windows and Linux. The Ryzen AI Max+ 395 puts 16 CPU cores and a 40-unit Radeon GPU on one package fed by a 256-bit LPDDR5X-8000 bus, and on a 128 GB machine you can assign up to 96 GB exclusively to the GPU, as HP’s ZBook Ultra page puts it. The same 75 percent rule as the Mac, arrived at from the other direction. Framework’s desktop version of the chip quotes the identical 96 GB ceiling.

A laptop with a separate graphics card is on the other ladder. The model must fit inside the card’s own VRAM, and NVIDIA rations it by tier: 8 GB on the RTX 5050, 5060, and base 5070; 12 GB on the 5070 Ti; 16 GB on the 5080; 24 GB on the 5090. That is the entire ladder. Above 24 GB there is no rung, at any price, on any laptop. The 64 GB of system RAM in a gaming laptop does not help, because a model spilling from VRAM to system RAM crosses a narrow PCIe link and slows to a crawl. AMD’s own discrete laptop line tops out at 16 GB, and machines carrying it are rare, so in practice the VRAM ladder is NVIDIA’s.

Put the ladders side by side, as the hero above does, and the crossover is obvious. An 8 GB card and a 16 GB Mac both hold an 8B model. A 12 GB card and a 24 GB Mac both hold a 14B. A 24 GB RTX 5090 laptop and a 32 GB Mac both hold a 32B. Then the VRAM ladder ends, and the unified ladder has three more rungs. The reflex to buy “the one with the RTX badge” is exactly right if you will only ever run 8B to 14B models, and exactly wrong if you have a 70B model in mind.

Bandwidth is the second axis, and discrete GPUs win it

Memory size decides what runs. Bandwidth decides how fast it talks. Generating text is a strange workload: for every single token, the chip reads the whole model out of memory, so tokens per second scale almost linearly with the width of the pipe. Apple’s own machine learning team showed it cleanly when the M5 arrived. Bandwidth went from 120 to 153 GB/s over the M4, and generation speed rose 19 to 27 percent, almost exactly the 28 percent the pipe grew.

Memory bandwidth the GPU can read, GB/s
Generating text reads the whole model out of memory for every token, so tokens per second track this number. Violet is unified memory, amber is discrete VRAM. Vendor peak figures; Strix Halo measures closer to 212 GB/s in practice.
Panther Lake iGPU
Intel thin laptops
154
Apple M5
MacBook Air, base Pro
153
Strix Halo
Ryzen AI Max+ laptops
256
Apple M5 Pro
MacBook Pro
307
RTX 5070 laptop
8 or 12 GB
384
Apple M5 Max
MacBook Pro
614
RTX 5070 Ti laptop
12 GB
672
RTX 5080 / 5090 laptop
16 / 24 GB
896

This is where the discrete card earns its keep. GDDR7 on a wide bus is simply faster than LPDDR5X shared with a CPU: NVIDIA quotes 896 GB/s for the RTX 5080 and 5090 laptop parts, against 614 GB/s on the M5 Max, 307 GB/s on the M5 Pro, and 153 GB/s on the base M5. Strix Halo’s 256-bit bus works out to 256 GB/s on paper, and a widely cited community benchmark set measures about 212 GB/s under load. So a 24 GB RTX 5090 laptop and a 32 GB MacBook Pro hold the same 32B model, but the RTX laptop reads it roughly three times faster than an M5 Pro would.

The published numbers line up with the arithmetic. On the standard 7B 4-bit test in the llama.cpp project’s Apple silicon table, an M5 generates about 32 tokens per second, an M5 Pro 66, and an M5 Max 120. A Strix Halo machine lands near 53 on the same test in the community set above. StorageReview measured a Razer Blade 16 with an RTX 5090 at 135 W running Llama 3 8B through Ollama at about 108 tokens per second, roughly half of the desktop 5090 in the same test. Different harnesses, so treat the figure below as a ballpark, but the ordering is robust.

Published generation speed on a 7–8B model, tokens per second
Different harnesses and slightly different models, so read these as ballparks. Anything above 20 is faster than most people read.
MacBook Air / Pro, M5
llama.cpp, 7B Q4_0
32
Strix Halo laptop
llama.cpp Vulkan, 7B Q4_0
53
MacBook Pro, M5 Pro
llama.cpp, 7B Q4_0
66
RTX 5090 laptop, 135 W
Ollama, Llama 3 8B
108
MacBook Pro, M5 Max
llama.cpp, 7B Q4_0
120

Two things to take from that chart. First, everything on it is faster than you read, so for an 8B model bandwidth is a luxury, not a requirement. Second, the gap matters at the top of the ladder. Divide bandwidth by model size and you get a ceiling: a 43 GB 70B model on a 256 GB/s Strix Halo tops out near 6 tokens per second, on a 614 GB/s M5 Max near 14. Both hold the model. Only one of them makes it pleasant. That is the trade unified-memory buyers make, and the M5 Max is the one laptop that makes it lightly.

What each tier actually runs

Model sizes cluster into bands, so each rung unlocks a distinct class of capability. The fits below use the same 4-bit sizes as our hardware calculator and leave room for context. Apple’s MLX team gives a useful calibration for the 24 GB tier: a 24 GB MacBook Pro runs an 8B model at full precision or a 30B mixture-of-experts at 4-bit while keeping the whole workload under 18 GB.

Model size to memory tier
4-bit weights at about 0.6 GB per billion parameters, plus room for context. Same sizes as our memory-math and hardware-calculator posts, so the tiers line up across guides.
8B
4.9 GB
8 GB VRAM or 16 GB unified
Private chat, drafting, summaries. Every laptop in this guide.
12–14B
~8–9 GB
12 GB VRAM or 24 GB unified
Noticeably better writing and code. RTX 5070 Ti, 24 GB Air.
27–32B
~17–20 GB
24 GB VRAM or 32–48 GB unified
The everyday ceiling for most people. RTX 5090 laptop, 32 GB Mac.
70B
43 GB
64 GB unified
No laptop graphics card. M5 Pro at 64 GB, Strix Halo at 64 GB.
120B MoE
~65 GB
96–128 GB unified
gpt-oss-120b class. M5 Max at 128 GB, Strix Halo at 128 GB.

The row to circle is 70B. It is the point where open models stop feeling small, and it is the first row with no entry in the VRAM column. A 64 GB MacBook Pro or a 64 GB Strix Halo laptop is the cheapest portable machine that runs it at all. The 120B row is the same story one rung up. A Strix Halo owner on Ollama’s issue tracker reports about 110 GB available to the GPU, with 30B to 35B models at 8-bit generating around 40 tokens per second, and the community benchmark set runs the 58 GB Llama 4 Scout at about 20. Those are workloads a 24 GB card cannot start.

The three shapes on sale: a thin unified-memory laptop, a Strix Halo tablet, and a thick machine with a graphics card. Which one is right depends on the model size you name, not the badge. Illustration generated with Seedream 5 Pro via Runware.

The lineup by platform

Mac. Summarized here, argued in full in the Mac guide. The MacBook Air configures to 16, 24, or 32 GB at 153 GB/s and starts at $1,299 after June’s increase. The MacBook Pro with M5 Pro goes to 64 GB at 307 GB/s from $2,499, and the M5 Max to 128 GB at 614 GB/s from $4,099 for the 14-inch. Metal support is built into every local runtime, including the Ollama that CSuite bundles, so there is no driver story to tell.

Windows, unified. Three Strix Halo laptops matter. The ASUS ROG Flow Z13 is a 1.2 kg 13-inch tablet with 32, 64, or 128 GB; the 128 GB model is on ASUS’s own store and US retailers listed it around $3,300 in early September. The HP ZBook Ultra G1a is the 14-inch workstation version, up to 128 GB, which launched at $4,296 for 32 GB and $8,250 fully loaded; street prices have fallen well below list since. The ASUS TUF Gaming A14 puts the newer Ryzen AI Max+ 392 in a $2,200 gaming laptop, but the US version stops at 32 GB, which is the wrong end of the unified ladder to pay a premium for. Buy Strix Halo at 64 or 128 GB or not at all.

One more entry is on the horizon rather than the shelf. AMD showed the successor, Ryzen AI Max 400 or “Gorgon Halo,” in July: up to 192 GB of LPDDR5X-8533 on the same 40-unit GPU, with systems due later in the second half of the year. A 16-inch HP ZBook Ultra carrying it has surfaced in benchmark listings but is not announced. If 128 GB is your target and you can wait a quarter, waiting is rational.

Windows, discrete. Sort by VRAM and ignore the rest. Dell’s catalog is a clean illustration of the ladder’s pricing: the Alienware 16 Aurora starts at $1,099.99 with up to an RTX 5070, the 16X Aurora with a 12 GB RTX 5070 Ti at $2,199.99, and the 16 and 18-inch Area-51 with the 24 GB RTX 5090 from $3,449.99 and $3,999.99. Lenovo, ASUS, MSI, and Razer sell the same chips in the same bands. The RTX 5070 exists in 8 and 12 GB versions; if a listing does not say which, assume 8 and ask. On Windows, CSuite’s bundled Ollama uses CUDA on these cards, and Vulkan on AMD and Intel graphics.

The “don’t” tier. Ordinary thin laptops with Intel or AMD integrated graphics share system memory the way a Mac does, but through a narrow pipe. Intel’s Panther Lake tops out at LPDDR5X-9600 on its fastest parts, roughly 150 GB/s, an M5-class pipe feeding a much smaller GPU with no Metal and no CUDA behind it. These machines run an 8B model at readable speed, which is a fine thing for a general-purpose laptop to do. They are not what to buy for local AI. The same goes for the NPU on any “AI PC” sticker: it serves small on-device models, not the 12B to 70B class this guide is about.

Sustained inference is where thin laptops lose

Copper heat pipes and a fan inside a laptop. How much of this a chassis carries decides the power its GPU is allowed to draw, which decides how fast the same chip runs. Photo by Vishnu Mohanan on Unsplash.

A laptop GPU name is not a spec. NVIDIA sells the RTX 5090 laptop chip to manufacturers with a graphics power range of 95 to 150 W, and each maker picks a point on that range that its chassis can cool. The thin machine and the thick machine carry the same badge and run at different speeds. Notebookcheck measured an ROG Zephyrus G16’s 120 W RTX 5090 at 25 to 30 percent slower than the 175 W version in a Schenker Neo 16. Same chip, same 24 GB, a quarter of the speed left on the table for thinness.

Inference is a sustained load, not a burst. A long document summary or a batch of image generations keeps the GPU pinned for minutes, and a chassis that cannot shed the heat throttles clocks and spins fans to the limit. If you are choosing between two laptops with the same GPU, buy the thicker one, or the one whose reviews quote the higher sustained power figure. Ask the spec sheet for the number, not the name.

Then plug it in. NVIDIA’s BatteryBoost turns on automatically the moment a laptop is unplugged and regulates the GPU down to stretch battery life; ASUS documents the same behavior as the default on its ROG machines. A discrete-GPU laptop on battery is a different, slower computer, and a GPU drawing 100 W or more empties any battery an airline will let you carry in about an hour. The unified-memory machines, the fanless MacBook Air above all, are the ones that can genuinely run a model on a train. That is a real advantage, separate from memory, and it belongs in the decision.

Linux runs the same laptops, with driver homework

There is no Linux-specific laptop hardware. There is a driver stack per vendor, and it decides how much of your first weekend goes to setup. NVIDIA is the low-friction path: Ollama supports any card with compute capability 5.0 or newer on driver 550 or newer, which covers every RTX laptop chip sold today, and Ubuntu 26.04 LTS is the first release to carry CUDA and ROCm in its own repositories, on a Linux 7.0 kernel. The laptop-specific wrinkle is hybrid graphics: the integrated GPU drives the screen and the NVIDIA chip sleeps until asked. NVIDIA’s driver handles this with PRIME render offload, and compute runtimes like Ollama find the card without it, so in practice it just works once the driver is installed.

On Linux the laptop choice is driver-shaped: NVIDIA needs a render offload setup on hybrid machines, Strix Halo wants a recent kernel, and Vulkan works on everything. Photo by Gabriel Heinzer on Unsplash.

Strix Halo on Linux is capable and young. The chip, gfx1151 in ROCm’s naming, sits outside ROCm’s official support list, and the recipes that work today lean on the HSA_OVERRIDE_GFX_VERSION=11.5.1 override and a fresh kernel. Kernel version is not a formality here: the community benchmark set found a 15 percent performance difference between Linux 6.14 and 6.15 on otherwise identical systems. Which backend is faster has flipped between releases, with Vulkan leading in the spring 2025 numbers and ROCm recommended in the 2026 Ollama guide, so check the current guidance for your Ollama version before assuming. Budget an evening, and expect it to get shorter with every kernel.

Vulkan is the universal fallback. Ollama enables it by default on Windows and supports it on Linux with the vendor’s Vulkan driver installed, which is how AMD and Intel integrated graphics get any acceleration at all. CSuite’s Linux build follows the same split: text models go through Ollama with CUDA on NVIDIA machines, while its local image and video runtime ships Vulkan-only on Linux, so an NVIDIA laptop generates images through Vulkan there. If you want the fuller picture of what runs where on this platform, start with the Linux apps guide.

The shortlist at three budgets

Prices are September 2026 US list prices from the vendor pages linked above, and they will drift, mostly upward, while the memory shortage lasts. The picks will not: they follow from the ladders, and the ladders are fixed by silicon.

One pick per platform at three budgets
About $1,5008B to 14B models, private chat and drafting
Mac
MacBook Air M5, 24 GB
Silent, runs a 14B or a 30B MoE at 4-bit. The best buy at this price.
Windows
RTX 5070 Ti laptop, 12 GB
Fastest 8B to 14B tokens at this price. Expect to stretch toward $2,000 for the 12 GB card; the $1,100 tier is 8 GB.
Linux
The same RTX 5070 Ti laptop
NVIDIA is the low-friction driver path. Skip the 8 GB cards.
About $2,500 to $3,00032B daily, 70B within reach
Mac
MacBook Pro 14″ M5 Pro, 48–64 GB
64 GB is the cheapest laptop that runs a 70B model at all.
Windows
ROG Flow Z13, 64 GB Strix Halo
Unified path. RTX 5090 laptops are faster but stop at 24 GB.
Linux
Strix Halo at 64 GB, recent kernel
Works well on a current kernel; budget an evening for setup.
$4,000 and up120B-class models, several models resident
Mac
MacBook Pro M5 Max, 128 GB
614 GB/s makes 70B pleasant and 120B-class usable.
Windows
HP ZBook Ultra G1a or Flow Z13, 128 GB
96 GB to the GPU. Or wait: 192 GB Gorgon Halo laptops are due.
Linux
Strix Halo at 128 GB
The only Linux laptop that holds gpt-oss-120b. Kernel 6.15 or newer.

The pattern across all three rows is the thesis of this guide. At $1,500, the VRAM ladder and the unified ladder are neck and neck, and the discrete card wins on speed if 14B is all you need. At $2,500, the VRAM ladder has run out of rungs and every pick is unified memory. At $4,000, the only question left is whether you want 614 GB/s on macOS or 96 GB of GPU memory on Windows or Linux, and both are good answers.

If you cannot name the model you would run on 128 GB, you do not need 128 GB. Name the model, find its row in the table above, and buy one tier over it. Then, whichever ladder you chose, put a GPT-4-class model on it the first evening. The whole point of buying the memory is to use it.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app