Should you buy a GPU or rent one?
At today’s card prices, owning only wins if the GPU works seven hours a day. The formula, with real numbers, so you can find your line.
On this page
In January 2025 NVIDIA put a $1,999 price on the RTX 5090. At the end of September 2026 the cheapest one a US price tracker could find was $6,995. The card did not get faster. Everything around it got more expensive, and that one fact rewrites the oldest argument in local AI: whether to buy a graphics card or rent one by the hour.
This post works the math for three kinds of user, with every number pulled from a live price page on October 7, 2026, and the formula left in the open so you can swap in your own electricity rate and your own hours. The short version: owning only wins if the card works most of the day, and for text through open models it never wins on tokens alone. The reasons to buy are real, but they are not on the invoice.
The card you wanted costs three and a half times its launch price
NVIDIA’s January 2025 announcement priced the RTX 5090 at $1,999, the RTX 5080 at $999 and the RTX 5070 Ti at $749. Tom’s Hardware keeps a running tracker of the lowest US price for every card. On September 28, 2026 it read $6,995 for the 5090, $1,638 for the 5080 and $1,119 for the 5070 Ti. Even the RTX 5060 Ti 16 GB, the sensible entry card, sat at $779 against a $429 list price. A separate tracker saw the 5090 at $7,449 on Amazon two days later, up 40% in a single month.
The cause is memory. TrendForce reported that conventional DRAM contract prices rose 93% to 98% in the first quarter of 2026 as cloud providers bought up supply for inference servers, and a graphics card is a slab of memory with a chip attached. Used cards followed. A six-year-old RTX 3090, launched at $1,499 in September 2020, averaged $1,400 on eBay on October 7, 2026, up from $1,000 in May. Six-year-old hardware is trading near its launch price.
Owning has four costs and only one is printed on the box
The purchase price is the number everyone argues about, and it is the least interesting one, because it is not a cost until you divide it by time. A card you keep for three years costs a thirty-sixth of its price every month, minus whatever you sell it for at the end. We use 36 months and a resale of zero throughout this post. That is pessimistic on purpose: if owning wins with those assumptions, it wins with any.
The second cost is electricity. NVIDIA rates the RTX 5090 at 575 W, and Tom’s Hardware measured 572 W on average in a demanding game, so the ceiling is real. The US residential average was 18.31¢ per kWh in July 2026, from 13.41¢ in North Dakota to 48¢ in Hawaii. At the national rate, a flat-out 5090 costs 10.5¢ an hour. Chat and image work rarely pin the card, so this is also a ceiling, but it is the honest number to plan with.
The third cost is the rest of the machine: a power supply NVIDIA says should be at least 850 W, a case with room for a card that spans three slots, and the system memory that is going through the same price spike. The fourth is the one nobody budgets: the idle hours. A desktop left on for a model that answers once an hour is still drawing power for the other fifty-nine minutes. We leave both out of the tables, which again favours owning.
(price − resale) ÷ months + hours × watts ÷ 1000 × ratehours × hourly rate + storage GB × monthly rateinput tokens × input rate + output tokens × output rateRenting the same card costs about a dollar an hour
The rental market sells the exact card you were going to buy. RunPod lists an RTX 5090 pod at $0.99 an hour on its Secure Cloud and $0.69 on its Community Cloud, where the machines belong to other people. An RTX 4090 is $0.74 or $0.34. A used-market favourite, the RTX 3090, is $0.50 or $0.22. Billing is per second, so an hour of thinking and ten minutes of generating costs ten minutes. A 100 GB storage volume to keep your models between sessions is $7 a month.
That is not one vendor’s quirk. A comparison of 23 providers on October 7 put the median on-demand RTX 5090 at $0.70 an hour, with the cheapest in-stock listing at $0.35. For the data-centre class, Lambda charges $3.29 an hour for an H100 and $1.99 for an A100, billed by the minute. A data-centre card for the price of a coffee an hour is a thing you could not buy at any price two years ago.
The catch is the one you would expect. Your prompt, your document and your model leave your machine and run on hardware you do not control. You also pay in minutes rather than dollars: a pod takes time to start, a 20 GB model takes time to pull, and a popular card can be out of stock in the region you want. For a two-hour session that overhead is noise. For a two-minute question it is the whole session.
Three usage profiles get three different answers
Ten hours a month is a weekend project: a batch of product photos, a podcast to transcribe, an evening of trying a new model. Sixty hours is a daily habit, two hours every day with a local assistant open beside your work. Seven hundred and twenty hours is a card that never stops, running an agent, a watch folder, or a small service for a team. The table runs each profile against three cards you might actually buy.
| Hours a month | RTX 5090, new(32 GB, $6,995) | RTX 3090, used(24 GB, $1,400) | RTX 5060 Ti, new(16 GB, $779) | |||
|---|---|---|---|---|---|---|
| Own | Rent | Own | Rent | Own | Rent | |
Occasional, 10 h a weekend project | $195 | $16.90 | $39.47 | $12.00 | $21.97 | $9.70 |
Daily, 60 h two hours a day | $201 | $66.40 | $42.40 | $37.00 | $23.62 | $23.20 |
Always on, 720 h an agent that never sleeps | $270 | $720 | $81.08 | $367 | $45.37 | $201 |
The occasional user should never buy. At ten hours a month, even the $779 card costs more than twice what renting a bigger one does, and the $6,995 card costs eleven times more. The daily user is on the line: a used 24 GB card loses to the rental by five dollars, the 16 GB card is a dead heat, and the new flagship loses by $134 a month. Only the always-on user has a clear case, and there the gap is wide, because 720 hours of anything at a dollar an hour is $720.
Notice what sorts the rows. It is not the card. It is the hours. The same RTX 5090 is the worst purchase on the page at ten hours and the best at 720. That is the whole post in one sentence: a GPU is a fixed cost, and fixed costs only make sense when you spread them thin.
The break-even moved from two hours a day to seven
Set the owning bill equal to the renting bill and solve for hours, and you get the point at which the card has paid for itself. At its launch price, the RTX 5090 crossed that line at 55 hours a month, less than two hours a day. Anyone who used it most evenings came out ahead. At $6,995 the line sits at 212 hours a month, about seven hours every day of the month for three years. Against the cheaper Community rate it is nearly eleven.
The smaller cards hold up better, because their prices rose less in dollars even where they rose as much in percent. A used RTX 3090 pays for itself at about two and a half hours a day against the $0.50 Secure rate, and the 5060 Ti at about two. If you are going to buy in this market, the arithmetic points at the cheap end: a 24 GB card from 2020 that holds the same models as the new ones, slower. Our VRAM-first build guide explains why memory, not speed, decides what you can run.
One more thing the chart shows. Owning is a bet on the price staying down and the card staying useful. Renting is a bet on the rental rate staying down. In October 2026 the second bet looks safer. The rental rate is set by 23 providers competing for your hour, and the median sits at $0.70. The card price is set by a memory shortage, and it more than tripled. Only one of those two numbers has a shortage behind it.
Per token, the API beats your card on electricity alone
The hours comparison assumes you need a whole GPU. For text, you often do not. Open-weight models are served by the token, and the rates are small. On October 7, OpenRouter listed Qwen3 32B at $0.08 per million input tokens and $0.28 per million output tokens, Gemma 4 31B at $0.09 and $0.34, and gpt-oss-120b at $0.037 and $0.17. Those are models that fit on the cards above.
Now run the card as hard as it will go. Hardware Corner measured an RTX 5090 generating 61 tokens a second with Qwen3 32B at 4-bit, which is about 220,000 tokens an hour if it never pauses. The API would charge 6 cents for those tokens. The card’s electricity for that hour, at its rated power, is 10 cents. Before you count the $6,995, before you count the power supply, the open-model API is already cheaper per token than the electricity in your own card.
Real use makes it worse for the card, not better, because nobody generates 220,000 tokens an hour. A heavy daily user might produce a few million output tokens a month. At Qwen3 32B’s rate that is a dollar or two. No card on this page gets its monthly cost under $20. The comparison only flips if you want a model nobody will sell you by the token, or if you want the frontier: GPT-5.5 is $5 in and $30 out per million, and no card runs it at all. We made the longer version of this argument in cost per task, not cost per token.
What the math leaves out is why most people still buy
If the tables were the whole story, almost nobody should own a GPU in October 2026, and yet the cards still sell at $7,000. The buyers are paying for things the formula cannot price. Privacy is the first. A card in your own machine means the contract, the medical note, the unreleased design never leaves the room, and no terms of service changes that next quarter. For a lawyer or a clinic that is not a preference, it is the job.
Offline is the second. We have made the case in the airplane-mode test: a tool that stops when the network stops is a tool you are borrowing. A rented pod is the most borrowed computer there is. The third is the zero-marginal-cost habit. When each extra attempt is free, you try the eleventh prompt, regenerate the image again, run the test suite one more time. People with a card on their desk use AI differently from people watching a meter, and that difference is the point of having one.
Resale is the fourth, and it cuts both ways. The 3090 holding 93% of its launch price after six years means today’s buyer may lose far less than our write-off assumes, which drags every break-even in this post toward buying. It also means the market could turn: the same memory shortage that lifted used prices will end, and a card bought at the top will fall with it. Treat resale as a hedge you hope for, not a line you plan on. And if the card goes into a machine you need anyway, the Mac-versus-NVIDIA question changes the arithmetic again, because unified memory is priced as a computer, not as a graphics card.
Rent until you can count the hours
Here is the rule the numbers support. If you cannot say how many hours a month a GPU would work for you, rent one for a month and find out. The bill is the answer. Under two hours a day, keep renting, or pay by the token for open models and spend the saved thousands on anything else. Between two and seven hours, the cheap cards are the only purchases that pencil out: the 16 GB card from about two hours a day, the used 24 GB card from about two and a half, and both narrowly. Past seven hours, or past the point where the data cannot leave the building, buy, and buy the memory, not the speed.
Then do the sum again next spring. Every input in this post is a price that moved this year, and the biggest one, the card, moved by a factor of three. The formula is the part that lasts. Plug in your electricity rate, your hours and whatever the tracker says that morning, and the line will tell you which side you are on. Our cost of one hour of AI post prices the cloud side of that hour in more detail, and BYOK vs. SaaS AI covers what you own in each case. CSuite runs the same models on your card or through a provider, so you can switch sides when the line moves without switching tools.

