"Local is free" is the most common false claim in this field, and "local is never cheaper" is the second. Both dissolve under arithmetic. Here it is, with numbers you can substitute your own into.
Electricity per million tokens
Measure at the wall if you have a meter. Otherwise: GPU power draw plus 80–120 W for the rest of the machine.
Take a used 24 GB card at 350 W under load, plus 100 W system, so 0.45 kW. At $0.15/kWh that is $0.0675 per busy hour.
Now convert to tokens.
One user, 8B at Q4, 60 tok/s: 216,000 tokens/hour.
$0.0675 / 0.216M tokens = $0.31 per million output tokensBatched, 32 concurrent requests, ~1,600 tok/s aggregate: 5,760,000 tokens/hour.
$0.0675 / 5.76M tokens = $0.012 per million output tokensSame hardware, same electricity bill, 26 times cheaper per token. That single comparison is the whole economics of self-hosting: local is cheap when the GPU is busy, and expensive when it is not.
Hardware amortised
Say the card cost $700 and lasts three years.
- Busy 24/7: 26,280 hours, so $0.027/hour. Adds about $0.12 per million at single-stream, half a cent at batch 32.
- Two hours a day: 2,190 hours, so $0.32/hour. Adds about $1.48 per million at single-stream.
An idle GPU you have already bought produces the most expensive tokens you will ever generate.
Against API prices
Check current prices, because they move, but the shape holds. Hosted 8B-class open models sit around $0.05–$0.60 per million output tokens. Frontier models sit around $1.50–$15.
Break-even in tokens:
hardware_cost / (api_price_per_token - your_electricity_per_token)At a generous $0.60/M API price and $0.012/M electricity:
$700 / $0.588 per M = 1.19 billion output tokensAt batch throughput of 5.76M tokens/hour, that is 207 hours — about nine days of a saturated GPU. At two hours a day of one person chatting at 60 tok/s, the same 1.19 billion tokens takes roughly seven and a half years.
If the API alternative is a cheaper $0.10/M model, break-even moves to about 8 billion tokens — still under two months of a genuinely busy card, still never for casual chat.
Where you are changes the answer
Electricity is roughly $0.10/kWh across much of the United States and India, $0.25 in the UK, $0.30–0.40 in Germany and Denmark, and under $0.05 in parts of the Gulf. At $0.35/kWh the single-stream figure becomes $0.73 per million, and local loses outright on cost at low utilisation. Substitute your own rate; the sums above take thirty seconds to redo.
The costs people leave out
Your time, which usually exceeds everything else combined for light use. The machine you cannot use for anything else while a model is resident. Cooling, and noise if it sits in a room you work in. And the second card you buy eight months later.
The three honest verdicts
One person, chatting a few hours a day, small model. The API is cheaper, often by ten times or more. Run locally for privacy, offline capability, or control, and say that plainly instead of calling it a saving.
Sustained batch work — classifying five million documents, bulk rewriting, generating embeddings, an overnight pipeline — on hardware you own. Local can be five to fifty times cheaper, and the batching arithmetic above is exactly why.
Frontier-quality output. There is no local option at a price you would pay. A machine that runs the largest open models well costs more than a decade of most people's API bills, and depreciates while it sits there.
Work out your own break-even before you buy anything. It is one division.
Before you move on