ThinkFacility

Blog

We priced a home LLM rig against a $20 subscription. It breaks even in 23 years.

Every number below is from the labs' own pricing pages and our own price table, checked on September 5, 2026.

People buy hardware to run an LLM at home for lots of reasons. Saving money is the one that doesn't survive contact with a spreadsheet.

What it costs a month, everything on one scale

OptionPer month
Gemini AI PlusGoogle$4.99
ChatGPT GoOpenAI, may include ads$8.00
Local rig, year two onelectricity only, 2 h a day$11.00
Gemini AI ProGoogle$19.99
ChatGPT PlusOpenAI$20.00
Claude ProAnthropic, $17 on annual$20.00
ChatGPT ProOpenAI$100.00
Claude MaxAnthropic, from$100.00
Gemini AI UltraGoogle, 5x Pro limits$99.99
Local rig, first year$2,500 hardware spread over 12 months, plus power$219.00

Subscription prices read off each lab's own pricing page on September 5, 2026. The local rig is a used RTX 3090 build at $2,500, running two hours a day.

The rig saves you $9 a month, and cost $2,500.

That's 278 months to break even against a Claude Pro or ChatGPT Plus subscription. Call it 23 years, by which point the graphics card is a museum piece.

Where the $2,500 goes

PartCostNote
Used RTX 3090, 24GB$1,000–1,400Secondhand only. New ones stopped existing.
Processor~$500Ryzen 9 9950X, $499.99 against a $649 list
Power supply~$150Corsair RM850x, 850 W
Memory, 64GB DDR5$1,000–1,150Mid-shortage. See below.
Board, case, storage~$500Estimated, not individually priced
Electricity$11–20.50Per month, 2 to 8 hours a day

The two product links above are affiliate links, so we earn a commission if you buy through them. They're the only parts of this build worth pointing at today: both are sold by Amazon at sane prices. The graphics cards and memory kits I checked were mostly single-unit third-party listings at inflated prices, and I won't link those at any commission.

Two numbers for how strange this market is right now. A discontinued RTX 4090 that launched at $1,599 is listed at $3,799 by a seller with one unit left. And a 64GB memory kit went from $999.99 to $1,144.22 between two checks the same afternoon, which is why there's a range in that table instead of a price.

What you get for it

Local rig$20 subscription
Up front$2,500Nothing
Model you getOpen weights, about 32B, squeezed to 4 bitsThe lab's current frontier model
Speed~27 words a secondFaster, and it isn't your electricity
Your text leaves the buildingNeverEvery time
Works with the internet downYesNo
Usage capsNo capYes, and they move
Model can be discontinuedNo, it's on your diskYes, labs retire them on a date
Setup and maintenanceYoursNobody's

The one case where local wins on money

Per-token API pricing is a different question from subscriptions, and it's the only place the arithmetic flips.

Per 1M words outCostAgainst the rig
Your electricity$0.61325 W for 10.29 hours
Ministral 3 14B$0.20Cheapest we list. 3x cheaper than your power bill.
Claude Opus 5$25.0041x your electricity
GPT-6 Astra$50.0082x your electricity
GPT-5.5 Pro$180.00295x your electricity

API prices from our price table, which is rebuilt every fifteen minutes from each lab's own docs.

So it depends entirely on what you'd otherwise be paying for. Replacing a cheap hosted model, the electricity alone costs three times more and you never catch up. Replacing a frontier model at $50, a $2,500 rig pays for itself after about 51 million words, which is roughly 17 months at 100,000 a day.

That only holds if a 32B open model actually does your work. That's the real question, and it isn't a pricing question.

What people run

SetupSpeedCatch
No graphics card~4 w/sSlower than you read. Memory bandwidth is the wall, not the processor.
One 24GB cardRTX 3090 or 4090~27 w/sThe default. 32B models fit at 4-bit.
Two used 3090s~27 w/s48GB gets you a 70B model. Double the heat.
Mac Studio Ultrathe premium answer~10–15 w/sMemory is shared with the chip, so it holds models no graphics card can. Sips power: 9 W idle, 270 W flat out.

The insane end

If money isn't the constraint, the ceiling has moved and the prices have gone somewhere silly. Two of these cards cost more than the whole build in the table above.

PartTodayLaunched atWhat it buys
RTX 5090, 32GB$5,150$1,9992.6x its own launch price. In stock, which is the surprising part.
RTX 4090, 24GB$3,799$1,599Discontinued in 2024 and still going up. One left at that seller.
Mac Studio M5 Ultra, 96GB$5,499From Apple direct. Amazon doesn't stock the current ones.
Mac Studio M5 Max, 36GB$2,499The sane half of the same idea.
RTX 4060 Ti, 16GB$1,364$499A budget card at 2.7x list. One left.
64GB DDR5-6000 kit$1,144Was $999.99 earlier the same afternoon.
Mac mini M4, 16GB$1,050Apple charges $799 for the same machine. Buy it from Apple.

Product links are affiliate links, so we earn a commission if you buy through them. Prices read off Amazon and apple.com on September 5, 2026, in the middle of a memory shortage, and several of these are single-unit third-party listings that will not hold. We're linking them so you can see the numbers, not because we think you should pay them.

What the beast actually buys you

Here's the case for spending $10,000 on a Mac, and it isn't speed. It's that unified memory lets one machine hold a model with as many parameters as a frontier model. People have run 671-billion-parameter open models entirely in a maxed Mac Studio's memory, at somewhere around 10 to 15 words a second. No graphics card at any price does that, because no consumer card has that much memory on it.

What it doesn't buy you is a frontier model. It buys you a model of frontier size, which is a different claim. Whether it writes as well as GPT-6 Astra or Claude Fable 5.1 is exactly the thing we don't measure and won't guess at, and the open models are generally behind.

The money still doesn't work out. A $10,000 machine against a $20 subscription needs the subscription to be worth something it isn't. Against a $50-per-million frontier API it repays after about 200 million words, which is real if you're running a business on it and fantasy if you're not.

Buy the beast because you want the thing on your desk, or because your work genuinely can't leave the building. Both are good reasons. Neither is a saving.

What to install

SoftwareLicenseUse it if
OllamaMIT, freeYou just want a model running. The usual answer.
LM StudioClosed, freeYou'd rather click than type.
llama.cppMIT, freeYou want to tune it. Everything above is built on this.
vLLMApache 2.0, freeYou're serving a team, not yourself. Much faster at that.

So who should

Nobody doing it to save money on a small model. The top table is the whole answer and it doesn't improve with scale.

The good reasons are in the third table, and none of them are prices. Your text never leaves the building, which for some work is the requirement rather than a bonus. It runs with the internet down, nobody throttles you, and no lab can retire your model in six months. Plenty of people also do it because owning the thing is satisfying, which is a real reason that no spreadsheet will ever agree with.

New to this? The beginner's guide starts further back.