Addaly is in open beta. Things will change, and AI answers can be wrong — check anything that matters.

The honest cost comparison

Running Models Yourself · lesson 10 of 10 · 10 min

"Local is free" is the most common false claim in this field, and "local is never cheaper" is the second. Both dissolve under arithmetic. Here it is, with numbers you can substitute your own into.

Electricity per million tokens

Measure at the wall if you have a meter. Otherwise: GPU power draw plus 80–120 W for the rest of the machine.

Take a used 24 GB card at 350 W under load, plus 100 W system, so 0.45 kW. At $0.15/kWh that is $0.0675 per busy hour.

Now convert to tokens.

One user, 8B at Q4, 60 tok/s: 216,000 tokens/hour.

$0.0675 / 0.216M tokens = $0.31 per million output tokens

Batched, 32 concurrent requests, ~1,600 tok/s aggregate: 5,760,000 tokens/hour.

$0.0675 / 5.76M tokens = $0.012 per million output tokens

Same hardware, same electricity bill, 26 times cheaper per token. That single comparison is the whole economics of self-hosting: local is cheap when the GPU is busy, and expensive when it is not.

Hardware amortised

Say the card cost $700 and lasts three years.

  • Busy 24/7: 26,280 hours, so $0.027/hour. Adds about $0.12 per million at single-stream, half a cent at batch 32.
  • Two hours a day: 2,190 hours, so $0.32/hour. Adds about $1.48 per million at single-stream.

An idle GPU you have already bought produces the most expensive tokens you will ever generate.

Against API prices

Check current prices, because they move, but the shape holds. Hosted 8B-class open models sit around $0.05–$0.60 per million output tokens. Frontier models sit around $1.50–$15.

Break-even in tokens:

hardware_cost / (api_price_per_token - your_electricity_per_token)

At a generous $0.60/M API price and $0.012/M electricity:

$700 / $0.588 per M = 1.19 billion output tokens

At batch throughput of 5.76M tokens/hour, that is 207 hours — about nine days of a saturated GPU. At two hours a day of one person chatting at 60 tok/s, the same 1.19 billion tokens takes roughly seven and a half years.

If the API alternative is a cheaper $0.10/M model, break-even moves to about 8 billion tokens — still under two months of a genuinely busy card, still never for casual chat.

Where you are changes the answer

Electricity is roughly $0.10/kWh across much of the United States and India, $0.25 in the UK, $0.30–0.40 in Germany and Denmark, and under $0.05 in parts of the Gulf. At $0.35/kWh the single-stream figure becomes $0.73 per million, and local loses outright on cost at low utilisation. Substitute your own rate; the sums above take thirty seconds to redo.

The costs people leave out

Your time, which usually exceeds everything else combined for light use. The machine you cannot use for anything else while a model is resident. Cooling, and noise if it sits in a room you work in. And the second card you buy eight months later.

The three honest verdicts

One person, chatting a few hours a day, small model. The API is cheaper, often by ten times or more. Run locally for privacy, offline capability, or control, and say that plainly instead of calling it a saving.

Sustained batch work — classifying five million documents, bulk rewriting, generating embeddings, an overnight pipeline — on hardware you own. Local can be five to fifty times cheaper, and the batching arithmetic above is exactly why.

Frontier-quality output. There is no local option at a price you would pay. A machine that runs the largest open models well costs more than a decade of most people's API bills, and depreciates while it sits there.

Work out your own break-even before you buy anything. It is one division.

Before you move on

Two teams each buy the same $700 GPU to replace the same API. Team A runs a nightly classification job that saturates the card for six hours. Team B has five people chatting through the day, keeping the GPU busy about forty minutes in total. Who saves money?

Pick the one you would defend. Nobody sees your answer.

No ads. No data sale. No public scores on people. Ever.

© 2026 Addaly