Skip to content
CSuite
Oct 5, 20269 min read

Running local AI on every size of machine

A $45 board, a phone, a laptop and a workstation can all run AI offline. Which size of model fits yours?

On this page
Six sizes of computer · the model each one holds · October 2026
The smallest model here is 20 kilobytes. The largest is 142 gigabytes. Both run with the Wi-Fi off.
Microcontroller
“Yes / no” keyword spotter
20 KB
Raspberry Pi 5
Gemma 4 E2B
2.6 GB
Phone
Gemma 4 E4B
3.7 GB
Laptop
Gemma 4 12B
8 GB
Desktop with a GPU
Gemma 4 31B
20 GB
Workstation
Qwen3 235B
142 GB
Download size of one example model per machine. The bars use a log scale, so each step to the right is ten times larger. Amber fits in a pocket, violet sits on a desk. Sources are linked in the post.

Somewhere in your home there is probably a chip the size of a fingernail that listens for one word. It has less memory than a single photo takes up, it has never seen the internet, and it runs a small AI model all day. At the other end of the scale, a desktop the size of a shoebox can hold a model with 235 billion parameters. The file for the second model is about seven million times larger than the first.

People ask “can my computer run AI?” as if the answer were yes or no. The answer is almost always yes. The real question is which size of AI, and that depends on one thing more than any other: how much memory the machine has.

This post walks up a ladder of six machines, from a microcontroller to a workstation. Each rung answers the same three questions: what fits, how fast it runs, and what it is good for. The short version: find your machine in the table below, run the model on its row, and expect each step up to buy you a larger model rather than a faster small one.

One rule holds on every machine: the model has to fit in memory

An AI model is a file full of numbers. To run it, the computer loads that file into memory and reads through most of it for every word it writes. If the file is bigger than the memory, the model does not load, or it runs so slowly that nobody would wait for it.

The trick that makes models fit is quantization, which stores each number with fewer bits. Google’s own table for Gemma 4 shows the effect. Its 12B model needs 26.7 GB at full 16-bit precision and 6.7 GB at 4-bit. That is the same model at a quarter of the size, with a small loss in quality. Nearly every model named in this post is a 4-bit version. Our memory-math explainer has the rule of thumb, which is about 0.6 GB per billion parameters.

The ladder this post climbs, left to right. Every one of these can run AI with no internet connection. What changes is the size of the model. Illustration generated with Seedream 5 Pro via Runware.
MachineMemoryWhat fitsExample modelGood for
Microcontroller256–512 KBKilobytes. No language models.Keyword spotter, under 20 KBWake words, gestures, sensor alerts
Raspberry Pi 51–16 GB1–2 billion parametersGemma 4 E2B, 2.6 GBOffline voice projects, tinkering
Phone1–3 GB free for a model2–4 billion parametersGemma 4 E4B, 3.7 GBSummaries, rewrites, photo questions
Laptop16–32 GB4–20 billion parametersGemma 4 12B, 8 GBDaily chat, writing, transcription
Desktop with a GPU16–32 GB on the card20–31 billion parametersGemma 4 31B, 20 GBFast chat, coding help, images
Workstation96–512 GB120 billion and upgpt-oss-120b, 65 GBThe largest open models
Model sizes are download sizes at 4-bit, read October 5, 2026 from Google’s LiteRT model pages and the Ollama library. The phone row is the memory Google says its mobile builds need, not the phone’s total. Each section below links its sources.

One note on scope before the climb. CSuite, the app we make, runs on laptops and desktops with macOS, Windows or Linux. The first three rungs use other tools, and we name them as we go.

A microcontroller runs AI, but not the kind you chat with

A microcontroller is a whole computer on one chip, the kind inside a thermostat or a toy. Its memory is counted in kilobytes. The Arduino Nano 33 BLE Sense Rev2, a popular board for learning this, has 256 KB of working memory and 1 MB of storage. Espressif’s ESP32-S3 has 512 KB built in.

What fits: models measured in kilobytes. Google’s LiteRT for Microcontrollers (the current name for TensorFlow Lite Micro) fits its core runtime in 16 KB. Its starter speech example is a model under 20 KB that tells “yes” from “no.”

How fast: fast enough to keep up with a microphone. Espressif’s published benchmark for its WakeNet9 wake-word model on the ESP32-S3 lists 16 KB of internal memory plus 324 KB of add-on memory, and 3 milliseconds of work for each 32-millisecond slice of sound.

Good for: wake words, gesture detection, and sensors that notice when a machine sounds wrong. This field is called TinyML, and Arduino points beginners to Edge Impulse for training models. Not worth it here: any chatbot. The smallest language model in this post is about 10,000 times larger than the Arduino’s memory.

A classic Arduino Nano. The black square in the middle is the whole computer: processor, memory and storage on one chip. Photo by Nizzah Khusnunnisa on Unsplash.

A Raspberry Pi runs a real language model, slowly

A single-board computer is a different animal. The Raspberry Pi 5 has a four-core 2.4 GHz processor, runs Linux, and comes with 1 GB to 16 GB of memory. Memory is also why the price list looks odd in 2026. The 1 GB board is still $45, while the 16 GB board launched at $120 and now lists at $305 after three memory-driven price rises.

What fits: the smallest real language models, at 1 to 2 billion parameters. Gemma 4 E2B is a 2.6 GB file and used about 1.5 GB of memory in Google’s test on a Pi 5.

How fast: 7.6 tokens per second, after a 7.8-second wait for the first word. That is a steady trickle of words, usable for short answers. The next model up, Gemma 4 E4B, drops to 3.2 tokens per second, which tests your patience.

An add-on does not change the size class. The Raspberry Pi AI HAT+ 2 carries its own AI chip and 8 GB of memory. It launched in January at $130 and is listed at $200 today. The models offered for it at launch were all 1 to 1.5 billion parameters. Raspberry Pi is direct about the gap: it says edge models run at 1 to 7 billion parameters while the large cloud models range from 500 billion to 2 trillion.

Good for: an offline voice assistant, a smart-home brain, or learning how this works on a board you can hold in one hand. Speech fits well: whisper.cpp lists its tiny transcription model at about 273 MB of memory. Not worth it here: long documents, or anything above about 4 billion parameters.

A Raspberry Pi 4, the board before the Pi 5 and the same size. Unlike a microcontroller it runs a full operating system, which is what lets it load a language model. Photo by Jainath Ponnala on Unsplash.

Your phone is a faster AI computer than a Raspberry Pi

This is the rung that surprises people. A recent phone has a graphics chip, and that chip runs a small model about seven times faster than a Pi’s processor does.

One small model on seven machines, tokens per second
Gemma 4 E2B, the same 2.6 GB file everywhere, writing speed. A token is a short word or a piece of a longer one.
Raspberry Pi 5, 16 GB
processor only
7.6
Jetson Orin Nano
small GPU board
24.2
Intel Lunar Lake laptop
integrated graphics
48.4
Samsung Galaxy S26 Ultra
phone GPU
52.1
iPhone 17 Pro
phone GPU
56.5
Desktop, RTX 4090
24 GB graphics card
143.4
MacBook Pro, M4 Max
Apple silicon GPU
160.2
Google’s own measurements with its LiteRT-LM runtime, from the Gemma 4 E2B model page, read October 5, 2026. We did not run these ourselves. Amber rows are boards and phones, violet rows are laptops and desktops.

What fits: models of 2 to 4 billion parameters. Google lists its phone build of Gemma 4 E2B at 1.1 GB of memory and E4B at 2.5 GB. A phone shares its memory with every open app, so that is about the limit.

How fast: 52 tokens per second on a Galaxy S26 Ultra and 56 on an iPhone 17 Pro for E2B, with the first word in 0.3 seconds. The larger E4B runs at 22 to 25, still quick enough for chat.

You may already be using this rung. Apple lists Apple Intelligence on the iPhone 15 Pro and every iPhone 16 or newer. On Android, Gemini Nano runs inside a system service and gives apps on-device summaries, proofreading, rewrites and image descriptions. To pick the model yourself, the free Google AI Edge Gallery app runs Gemma 4 on Android 12 and iOS 17 or later, with no internet needed.

Good for: summarizing a message thread, rewriting a note, asking about a photo, and working on a plane or a train. Not worth it here: hard reasoning. A 2B model is quick and often wrong on multi-step problems. Our piece on why small models matter covers what they do well.

A phone has a graphics chip and gigabytes of memory, which makes it a better AI computer than most hobby boards. Illustration generated with Seedream 5 Pro via Runware.

A laptop is where local AI becomes a daily tool

A laptop is the first rung with enough memory for a model you would choose over a cloud chatbot for everyday work. It is also where CSuite starts. Laptops come in three kinds, and they behave differently.

  • Integrated graphics. A thin Windows laptop with an Intel Lunar Lake chip ran Gemma 4 E2B at 48 tokens per second and E4B at 25 in Google’s tests. That is phone speed, with more room.
  • Apple silicon. A Mac shares one pool of memory between the processor and the graphics chip. The MacBook Air starts at 16 GB and goes to 32 GB. The llama.cpp project’s gpt-oss guide lists a 36 GB M4 Max writing 92 tokens per second with the 20B model.
  • A separate graphics card. Here only the card’s own memory counts, and laptop cards carry less of it than desktop ones. Our laptop guide goes through the options.

What fits: on 16 GB, Gemma 4 12B at about 8 GB is the comfortable pick, and it reads images as well as text. On 32 GB, gpt-oss-20b at 14 GB fits with room for a long conversation. Our Gemma 4 guide has the commands.

Good for: chat, drafting, summarizing documents, and transcription. Whisper’s largest model needs about 3.9 GB, so it sits beside a chat model on a 16 GB machine. Not worth it here: the 30B class on a 16 GB laptop. It will not load, and no setting changes that.

On a desktop, the graphics card’s memory is the ceiling

A desktop adds a full-size graphics card, and the card brings its own fast memory. That memory is the number to shop by. Today’s top consumer card, the GeForce RTX 5090, has 32 GB. The RTX 5080 below it has 16 GB.

What fits: on a 16 GB card, gpt-oss-20b, which needs 14.9 GB with a short conversation. On 24 or 32 GB, Gemma 4 31B, which Google lists at 17.5 GB to load.

How fast: very. The same llama.cpp guide lists gpt-oss-20b at 282 tokens per second on an RTX 5090, 205 on a 16 GB RTX 5080, and 162 on a used RTX 3090. Answers arrive faster than you can scroll.

Good for: coding help, long documents, and image generation, which is built for these cards first. Not worth it here: the 120B class. It needs 64 GB, and no consumer card has that. Our Apple silicon vs NVIDIA comparison covers what happens at that wall.

A workstation buys room for the largest open models

The top rung is defined by memory counted in the hundreds of gigabytes. There are two common routes. NVIDIA’s RTX PRO 6000 workstation card has 96 GB. Apple’s Mac Studio with M5 Ultra starts at $5,499 and can be ordered with up to 512 GB.

What fits: gpt-oss-120b at 65 GB, and on the largest Macs, models like Qwen3 235B at 142 GB.

How fast: the llama.cpp guide lists gpt-oss-120b at 196 tokens per second on the RTX PRO 6000 and 80 on a 192 GB M2 Ultra Mac. Both are faster than a small model on a phone.

Good for: the strongest open models, several people sharing one machine, and work that cannot leave the building. Not worth it here: buying one for everyday chat. A 12B model on a laptop already handles that.

A bigger machine buys a bigger model, not a faster small one

Look at the speed chart again. The same small model runs at 56 tokens per second on a phone and 143 on a high-end desktop card. That is under three times the speed from a machine many times the size. Nobody reads at 143 tokens per second, so the extra speed is mostly wasted.

The value of each rung is the model it unlocks. A Pi holds a 2B model. A laptop holds a 12B model that makes fewer mistakes. A desktop card holds a 31B model, and a workstation holds one four to eight times larger again.

So the practical advice is short. Find your machine in the table and run the model on its row tonight. If the answers are good enough, stop there. If they are not, the fix is more memory, and our hardware calculator shows how much the next step needs.

More reading

One-time payment. Yours forever.

No subscriptions. No seats. No renewals. Buy CSuite once, future updates included.

Secure checkout via Stripe. Already have a license? Download the app