Skip to content
CSuite
Sep 30, 20269 min read

Apple silicon vs NVIDIA for local AI in 2026

An RTX 5090 wins by 2.4 times on a small model and loses by 2.6 times on a big one. Which side is your work on?

On this page
Same software · two model sizes · the lead changes hands
NVIDIA is faster while the model fits. The Mac is faster once it does not.
A 7B model, 4 GB
fits on both
RTX 5090, 32 GB
290 tokens/s
M5 Max, 40-core GPU
120 tokens/s
A 120B model, 59 GB
fits on the Mac only
RTX 5090, spilling to the CPU
about 30 tokens/s
M3 Ultra Mac Studio
78 tokens/s
Text generation in llama.cpp. Left: Llama 2 7B at 4-bit, from the project’s community benchmark threads. Right: gpt-oss-120b, the RTX 5090 figure from the llama.cpp maintainers’ guide with part of the model held in system RAM, the Mac figure measured by Jeff Geerling in December 2025. Sources are linked in the post.

Two people each buy a computer for running AI at home. One buys a Mac Studio with 128 GB of memory. The other builds a PC around a GeForce RTX 5090. A month later the PC owner is making images faster and running video tools the Mac owner cannot install. The Mac owner is chatting with a model so large that the PC owner can only run it at less than half the speed.

Both of them bought the right machine for something. The trouble is that “Apple silicon or NVIDIA” sounds like one question and is really two. How large a model can the machine hold? And how fast does it run the models that fit? Apple wins the first. NVIDIA wins the second, and also wins on which software works at all for images and video.

This post sorts the common kinds of local AI work onto one side or the other, using measurements where the same model ran on both. The short version: if your work is language models bigger than about 32 GB, buy the Mac. If it is images, video, or long prompts on small and mid-size models, buy the NVIDIA card. For everyday chat with small models, either one is already faster than you can read.

How big and how fast are two separate questions

A model has to sit in memory the graphics chip can reach. If it does not fit, it does not run slower by a little. It either fails to load or crawls. So the first number that matters is how much of that memory you have. Our memory-math explainer has the arithmetic: about 0.6 GB per billion parameters at 4-bit, plus room for the conversation.

The second number is memory bandwidth, which is how quickly the chip can read that memory. Writing each word means reading most of the model, so bandwidth sets the speed limit.

The two platforms made opposite trades. An NVIDIA card keeps a small amount of very fast memory on the card itself. The RTX 5090 has 32 GB, and no GeForce card has more. A Mac shares one large pool between the processor and the graphics chip. Apple’s Mac Studio goes to 128 GB with the M5 Max and 256 GB with the M5 Ultra, at 614 GB/s and 1.2 TB/s of bandwidth. That is a lot of memory read at a good speed, against a little memory read at a great speed.

The whole comparison in one picture. The big tank holds far more but pours through a narrower pipe. The small tank pours faster and runs out sooner. Illustration generated with Seedream 5 Pro via Runware.

When the model fits, NVIDIA is faster, and long prompts widen it

The cleanest public comparison is llama.cpp’s own benchmark, where owners run the same small model with the same command and post the result. On Llama 2 7B at 4-bit, an RTX 5090 writes 290 tokens per second. Apple’s fastest laptop chip, the M5 Max, writes 120. The flagship card is 2.4 times faster at generating text.

One small model on eight machines, tokens per second
Llama 2 7B at 4-bit (Q4_0) in llama.cpp, flash attention off. The bar is writing speed. The number on the right is reading speed, the rate at which the machine takes in your prompt.
RTX 5090
32 GB VRAM
290
14,073
RTX 5080
16 GB VRAM
182
8,297
RTX 5070 Ti
16 GB VRAM
177
6,952
RTX 3090
24 GB VRAM
158
5,175
M5 Max, 40-core GPU
up to 128 GB unified
120
3,220
RTX 5060 Ti
16 GB VRAM
91
3,737
M5 Pro, 20-core GPU
up to 64 GB unified
66
1,621
DGX Spark
128 GB unified
57
3,062
Community-submitted runs from llama.cpp’s Apple silicon and NVIDIA CUDA benchmark threads, read September 30, 2026. Rows were submitted on different llama.cpp builds, so treat gaps under 10% as noise.

Look further down the chart and the picture is less one-sided. The M5 Max beats the RTX 5060 Ti, and lands within reach of a used RTX 3090. For plain chat, every machine here writes faster than anyone reads.

The bigger gap is in the right-hand column. Before a model writes anything, it has to read your prompt, and NVIDIA cards read far faster: 14,073 tokens per second on the RTX 5090 against 3,220 on the M5 Max, which is 4.4 times. On a short question you will never notice. On a long one you will. In the llama.cpp maintainers’ gpt-oss guide, from August 2025, a 32,768-token prompt took an RTX 5090 about 5 seconds to read. An M3 Ultra Mac Studio needed about 24 seconds and an M4 Max about 58.

That matters most for coding agents and document work, which resend tens of thousands of tokens on every step. Apple has been closing this gap: it says the M5 Ultra reads prompts up to 4 times faster than the M3 Ultra, and Ollama’s switch to Apple’s MLX framework in March roughly doubled generation speed in its own test, from 58 to 112 tokens per second. The Mac is catching up on speed. It has not caught up.

A GeForce RTX card in a tower. Its memory sits on the card, right next to the chip, which is why it is fast and why there is only so much of it. Photo by Christian Wiediger on Unsplash.

Past 32 GB, the Mac runs models a graphics card cannot hold

Now make the model bigger than the card. OpenAI’s open gpt-oss-120b is 59 GB at its standard 4-bit size, almost twice what an RTX 5090 holds. The software does not give up. It keeps what it can on the card and leaves the rest in ordinary system RAM, where the processor handles it far more slowly.

The llama.cpp guide puts an RTX 5090 running this way at about 30 tokens per second with an empty context. A reader with an RTX 3090 and 64 GB of DDR5 measured 27. On a Mac Studio with an M3 Ultra, where the whole model sits in unified memory, Jeff Geerling measured 78 tokens per second. The machine that lost by 2.4 times on the small model wins by about 2.6 times on the large one.

A Mac Studio. The memory inside is one pool shared by the processor and the graphics chip, and the current model can be ordered with up to 256 GB of it. Photo by Peng Originals on Unsplash.

Keep going and the card drops out entirely. The same Mac ran Qwen3 235B, a 132 GB file, at 26 tokens per second, and the 377 GB DeepSeek R1 at 20. Nothing with a GeForce badge loads either one. The first fits a 256 GB Mac Studio today. The second needs the 512 GB option, which Apple says is coming in late October.

NVIDIA can match the capacity, at a price. The 96 GB RTX PRO 6000 Blackwell runs gpt-oss-120b at 196 tokens per second in the same guide, far ahead of any Mac, and costs about $20,000 for the card. And NVIDIA’s own small desktop, the DGX Spark, uses 128 GB of unified memory at 273 GB/s, the same design as a Mac. It behaves like one: room for huge models, and 57 tokens per second on the small benchmark, behind an M5 Pro. The trade comes from the memory design, whoever builds it.

Memory a model can use costs four times more on an RTX 5090

The memory shortage has bent this comparison out of shape. The RTX 5090 launched at $1,999. The lowest new listing we found on September 28 was $6,996, and the tracked price has swung between $4,200 and $7,370 since late July. A Mac Studio with an M5 Max and 128 GB costs $5,099, and that price includes the rest of the computer.

HardwarePriceMemory for the modelPer GBLargest class it holds
RTX 5060 Ti 16 GB$78016 GB$4912–14B text, most image models
RTX 5070 Ti$1,20016 GB$75the same models, twice as fast
RTX 5090$6,99632 GB$21932B text, large image and video
RTX PRO 6000 Blackwell$20,00096 GB$208120B-class text, fast
Mac mini, M5 Pro, 64 GB$2,699about 48 GB$5670B text at 4-bit
Mac Studio, M5 Max, 128 GB$5,099about 96 GB$53120B-class text
Mac Studio, M5 Ultra, 256 GB$9,499about 192 GB$49235B-class text
US prices read September 30, 2026. Mac prices are Apple’s store prices for the whole computer, with the GPU able to use about three quarters of unified memory. Card prices are the lowest new listings on videocardprices.com, gpuprix.com and Newegg, for the card alone. The PC around it is extra.

The fair way to compare is dollars per gigabyte the model can use. macOS lets the graphics chip take about three quarters of unified memory on machines above 32 GB, so a 128 GB Mac is a 96 GB graphics card for this purpose. On that basis the Mac Studio costs about $53 per gigabyte. The RTX 5090 costs $219, four times as much, before you buy a case and a power supply.

The chart has a second lesson. At the low end NVIDIA is the better deal. A 16 GB RTX 5060 Ti costs the same per gigabyte as the largest Mac and drops into a PC you may already own. If your models fit in 16 GB, the card is the cheaper route, and our PC build guide covers the rest of the box. The Mac’s price advantage starts where consumer cards stop.

Image and video tools are written for NVIDIA first

Language models are the Mac’s best case. Image and video generation are its worst, for two reasons. These models do heavy arithmetic on every step, so raw compute counts for more than memory size, and most of them fit on a 16 to 32 GB card anyway. And the people who release them test on NVIDIA hardware.

You can see it in the install instructions. HunyuanVideo’s requirements say “An NVIDIA GPU with CUDA support is required.” Wan 2.2 gives its hardware needs in NVIDIA cards and does not mention Macs. CUDA is NVIDIA’s own programming layer, and it does not exist on a Mac. Popular speed-up libraries such as FlashAttention depend on it too.

Image generation does run on Apple silicon. ComfyUI supports Macs through Metal, and Mac apps such as Draw Things are well tuned. It is slower, though, and we could not find an independent test that timed one image model with the same settings on both. What exists is vendor claims. NVIDIA said in September 2025 that FLUX.1 Krea ran 8 times faster on an RTX 5090 than on an M3 Ultra. Apple says the M5 Ultra is up to 4.3 times faster at text-to-image than that M3 Ultra. Read both with care, since each company chose its own test. Together they suggest a real gap that is shrinking.

NVIDIA keeps moving as well. Its March update cut memory use for FLUX.2 and LTX-2 video by 60% and made them 2.5 times faster with a number format built for RTX 50 cards. If pictures and clips are the main job, the card is the safe choice.

Two good answers to two different jobs: a small box with a lot of memory for large language models, and a tower with a fast card for images, video and long prompts. Illustration generated with Seedream 5 Pro via Runware.

The Mac saves power at idle, not per token

Apple publishes wall-socket figures for the Mac Studio: 7 W idle and 200 W flat out for the M5 Max, 9 W and 385 W for the M5 Ultra. An RTX 5090 is rated at 575 W for the card alone, and NVIDIA asks for a 1,000 W power supply. A Mac Studio is small and quiet enough to leave running on a shelf as a home AI server.

That does not make it more efficient per word. Geerling measured both kinds of machine on the same 7B model. The M3 Ultra Mac Studio drew 226 W for 104 tokens per second. A PC with an RTX 4090 drew 338 W for 183, using llama.cpp’s Vulkan backend. Divide it out and the two produce nearly the same tokens per watt, with the PC slightly ahead. The card uses more power and finishes sooner.

So the Mac’s saving is real in two places: the hours the machine sits idle, and the large models where a card would be dragging its work through system RAM. For a machine that works hard a few hours a day, the electricity bills come out closer than the wattage labels suggest.

Pick by workload, not by brand

Which side each kind of work falls on
Chat and writing with models up to about 14B
Either
Both are faster than you can read. Buy on price and the computer you already like.
Coding agents and long documents
NVIDIA
These send huge prompts over and over, and prompt reading is NVIDIA’s widest lead.
The biggest open models, 70B and up
Apple silicon
They fit in unified memory and do not fit on a consumer card.
Image generation
NVIDIA
It works on both, but the speed work and the newest tools land on RTX first.
Video generation
NVIDIA
Several official releases require CUDA outright.
A quiet machine that is always on
Apple silicon
A Mac Studio idles under 10 W and tops out below what one RTX 5090 draws.

Start by naming the largest model you want to run and what kind it is. If it is a language model above roughly 32 GB, the decision is made: you want unified memory, and a Mac is the most memory per dollar on sale. Our Mac buying guide walks the tiers.

If your work is images, video, or agents that chew through long prompts, the decision is also made. Buy the NVIDIA card with the most memory you can afford, and accept the 32 GB ceiling. On a laptop the same split applies with smaller numbers, as our laptop guide shows.

If you already own one and wonder what you are missing, the honest answer is the other half of the chart. Mac owners are missing speed on long prompts and first-day support for new image and video models. PC owners are missing the 70B-and-up models that need more memory than any consumer card has. Neither gap is a reason to switch unless that missing piece is the work you do most.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app