When self-hosting an LLM is actually cheaper than an API
The question is usually asked as if there were one answer. There is not, because self-hosting does not compete with 'the API'. It competes with a specific price per million tokens, and those prices differ by a factor of forty between a frontier model and a hosted open-weight one.
By Evdokiia Petrovskaia, Director, Dubir GroupPrices checked
Short version: if you are replacing a frontier closed model, a single GPU pays for itself somewhere around 19 million tokens a month. If you are replacing a hosted open-weight endpoint, it essentially never does. Decide which of those two you are actually doing before pricing hardware.
What self-hosting costs per month
Four lines, and only the first is obvious.
- Depreciation is the purchase price spread over the card's useful life, so a $2,900 RTX 4090 over 36 months is about $81 a month. The card you already own is not free. Counting it that way is the most common error in these comparisons.
- A 4090 draws 450 W flat out, and a real workload averages far less. At a 35% duty cycle and €0.28 per kWh, the electricity is about $32 a month.
- Then someone has to patch it, watch it, restart it and upgrade a runtime that ships breaking changes monthly. That someone is paid. Charging 20% of hardware plus power for it is deliberately conservative, and it is still $23 a month.
- Idle time is not a line item at all, which is why it gets missed. The hardware bill is identical whether the GPU is saturated or asleep, while the API bill goes to zero on a quiet week.
Total for the 4090 example: about $135 a month, of which $55 is recurring once the card is written off.
The breakeven depends on what you are replacing
Now the comparison. The same $135 a month, against different APIs at 20 million input and 5 million output tokens per month:
| Replacing | API cost/month | Breakeven | Payback |
|---|---|---|---|
| Frontier closed model | $180 | 19M tokens/mo | 23 months |
| Mid-tier closed model | $90 | 38M tokens/mo | 82 months |
| Qwen3.6 35B-A3B, hosted | $8 | 420M tokens/mo | never |
That last row is the one worth sitting with. Hosted open-weight inference is cheap enough that buying a graphics card to avoid an eight-dollar bill takes 420 million tokens a month to justify. At that volume you are running a serving operation. That is a different project.
Bigger hardware does not fix it
A $5,000 RTX 5090 at 50 million input and 12 million output tokens a month costs about $216 self-hosted against $220 on a mid-tier API. That is 35 months to payback, breakeven around 61 million tokens a month, and the card is close to the end of a three-year write-off by the time it has earned itself back. Bigger hardware raises the cost and the breakeven in step. The ratio is set by the API price you are escaping, not by the card.
Capacity is not the constraint
One card generating 300 tokens a second around the clock produces about 788 million output tokens a month. At 60 tokens a second, a dense model on a Mac, it is still 158 million. For almost every internal tool, one GPU has more throughput than the organisation has demand. You are buying idle capacity. Which is why the honest question at the end of a procurement meeting is not how fast the card generates, but how many hours a day anyone will be waiting on it.
The reasons that actually hold up
Cost is the weakest argument for self-hosting, and the one most often given. The durable reasons are the ones nobody puts in a spreadsheet:
- The data cannot leave. Regulated industries, client confidentiality, personal data that a processor agreement will not cover. This is why we deploy models on client hardware, and it does not depend on token volume at all.
- A fixed monthly number beats a variable one for some finance departments, even when the variable one is lower on average.
- Models you host are not deprecated, repriced or rate-limited without your consent.
- A local model answers without a round trip, and keeps answering when the connection does not.
Work out your own number
Put your actual monthly token volume, your actual electricity price and the actual card you are considering into the self-hosted vs API calculator. Every figure above came out of it. If you have not yet decided which card, 4090 vs 5090 vs Mac compares the three on the two specifications that decide it.
Questions, answered
At what volume does self-hosting an LLM become cheaper?
There is no single number, because it depends on the API price you are comparing against. Replacing a frontier closed model, an RTX 4090 breaks even at roughly 19 million tokens a month. Replacing a mid-tier closed model, about 38 million. Replacing a cheap hosted open-weight endpoint at $0.15 per million input tokens, roughly 420 million a month, which is more than most companies will ever send.
How long does a GPU take to pay for itself?
In our worked example a $2,900 RTX 4090 replacing a frontier API at 25 million tokens a month pays back in about 23 months. Against a mid-tier model at the same volume, about 82 months, which is longer than the hardware's useful life. Against hosted open-weight models it does not pay back at all.
What costs am I forgetting when I self-host?
Electricity, which is small but not zero: a 4090 at 35% duty cycle costs about $32 a month at €0.28/kWh. Depreciation, which is the largest line and which people leave out entirely by treating the purchase as sunk. And operations: someone has to patch, monitor and restart it, and 20% of the hardware and power cost is a conservative allowance for that.
Does self-hosting mean I get the same quality?
Not automatically. The honest comparison is against hosting the same open-weight model through an API, and there the economics rarely favour buying hardware. The comparison that does favour it is against a closed frontier model, and that is a quality trade as well as a cost one. You are asking whether an open model is good enough for your task.
How many tokens can one GPU actually produce?
More than most people need. At 300 tokens a second running continuously, one card produces about 788 million output tokens a month. Capacity is rarely the binding constraint for an internal tool; the constraint is that you are paying for the card whether or not it is busy.
So when is self-hosting the right call?
When the data is not allowed to leave your infrastructure, when you need predictable cost rather than lowest cost, when latency to your own hardware matters, or when your volume against a frontier API is genuinely large. Privacy and regulation are better reasons than price, and they are the ones that hold up.
Private AI, from the spec to the running box
Dubir Group deploys open models on client hardware from Paphos, Cyprus, for the cases where the data is not allowed to reach a third-party API. We size the machine, install it and keep it answering.