Self-hosted LLM vs API: how to think about cost and break-even

APIs charge per token, so cost grows with usage but idle time is free. Self-hosting charges per GPU hour whether busy or idle, plus engineering time. Self-hosting wins only with high, steady utilization of a model that meets your quality bar; otherwise the API is usually cheaper.
What are you actually comparing?
An API bill is variable. You pay for input and output tokens, sometimes with discounts for cached prompts or batch jobs, and you pay nothing when traffic stops.
A self-hosted bill is mostly fixed. You rent or buy GPUs, pay for them every hour they exist, and add the people who keep the serving stack healthy. Traffic changes how much value you squeeze out of that fixed cost, not the cost itself.
So the real question is not “which is cheaper per token?” It is “how busy can I keep a GPU running a model that is good enough, and what does it cost to run that well?”
Quality is part of cost. If the model you can afford to self-host solves fewer tasks correctly, you pay in retries, human review and churn. A fair comparison fixes the quality bar first, then asks which option reaches it more cheaply.
How do the cost structures compare?
| Cost element | Hosted API | Self-hosted open model |
|---|---|---|
| Main unit | Tokens in and out | GPU hours, owned or rented |
| Idle cost | None | Full price of reserved capacity |
| Scaling up | Automatic within rate limits | Provision more GPUs, plan capacity |
| Engineering | Integration and prompt work | Serving, monitoring, upgrades, on-call |
| Hidden costs | Long prompts, retries, agent loops | Low utilization, failed nodes, storage, egress |
| Discounts | Prompt caching, batch processing, committed use | Reserved or spot instances, quantization |
| Quality ceiling | Frontier models available | Limited to open weights you can serve |
How do you calculate a break-even point?
Start from your own numbers, never someone else’s. Measure monthly tokens in and out from logs or a realistic forecast, then multiply by your provider’s current published prices, including any caching or batch discounts you would actually use.
For self-hosting, benchmark your chosen model on your chosen engine and hardware with real prompts. From that you get sustainable tokens per second at acceptable latency. Divide monthly tokens by that throughput to find GPU hours needed, then add headroom for peaks and redundancy.
Multiply GPU hours by the hourly price of the instances, add storage, networking and monitoring, and add a fair share of engineer time. The break-even is the volume where both totals meet. If you need redundancy across two machines for uptime, count both.
Run the calculation for three scenarios: today’s volume, a realistic twelve-month forecast and a pessimistic case where growth stalls. Self-hosting often looks great in the optimistic case and painful in the pessimistic one, and that asymmetry is a real risk.
Keep the spreadsheet alive. Re-run it whenever a provider changes prices, a new model changes quality per dollar or your traffic shape shifts, because each of these moves the break-even point.
Why is utilization the number that decides everything?
A GPU that is busy a small fraction of the day makes every token very expensive. The same GPU kept busy most hours, with batching engines such as vLLM packing many requests together, can make each token cheap.
Traffic shape matters as much as volume. Business-hours chat with quiet nights wastes reserved capacity, while overnight batch jobs such as document processing or embedding generation fill it. Mixing interactive and batch workloads on the same hardware is one of the most effective levers.
Autoscaling GPU servers helps, but cold starts for large models are slow, so few teams scale all the way to zero for latency-sensitive traffic.
Model size interacts with utilization. A model that fits on one GPU is far cheaper to keep busy than one that needs several cards with tensor parallelism, so choosing the smallest model that passes your evaluation is often the biggest single saving.
Which should you choose?
A simple rule of thumb without invented numbers: if you cannot state your measured tokens per month and your measured throughput per GPU, you are not ready to decide in favor of self-hosting yet.
Gather those two numbers first. They turn a debate about ideology into a spreadsheet anyone on the team can check.
- Low or unpredictable volume, early product, small team: API.
- You need frontier-level quality for the core feature: API, since the best models are not self-hostable.
- High, steady volume on a narrow task a smaller open model handles well: self-host, or use a provider that serves open models.
- Large overnight batch workloads: consider API batch discounts first, then self-hosting if you already run GPUs.
- Strict data residency or on-premise deals: self-host, and treat cost as secondary to winning the contract.
- Uncertain: stay on APIs behind a gateway and revisit when logs show stable, heavy usage.
Is there a middle option?
Yes. Many inference providers serve open-weight models through an API with per-token pricing. You keep open-model flexibility and portability without running GPUs, and can later move the same model in-house.
Dedicated or reserved endpoints from these providers sit between the two models: you pay for capacity, but someone else operates it. Compare these offerings on your own benchmark before committing.
Another middle path is splitting by workload. Keep interactive, quality-sensitive features on an API and move predictable bulk jobs, such as classification or embedding generation, onto self-hosted capacity that you can keep near full.
These hybrid paths also reduce lock-in risk, because the same open model can move between providers and your own hardware as prices and needs change.
How to cut cost on either path
Measure cost per feature and per customer, not just per month. That often reveals one expensive feature or one heavy customer driving most of the bill, which is far easier to fix than the whole system.
- Shorten prompts and cap output length; long system prompts repeated on every call add up.
- Use prompt caching for stable prefixes where your provider supports it.
- Route easy requests to a smaller model and reserve the large one for hard cases.
- Batch non-urgent work and run it off-peak or through batch endpoints.
- Quantize self-hosted models after checking quality on your evaluation set.
- Cache final answers for repeated questions.
Common mistakes in cost comparisons
- Comparing API list price with a GPU’s theoretical peak throughput instead of measured throughput at your latency target.
- Assuming 100 percent utilization for self-hosted capacity.
- Leaving engineering and on-call time out of the self-hosted column.
- Comparing a frontier API model with a much weaker open model and calling the saving real.
- Ignoring agent loops, retries and tool calls, which multiply token counts well beyond one prompt per user action.
- Using prices from an old blog post; both API and GPU prices move.
Where the self-hosting case breaks
It breaks when traffic is spiky, when the team is small, or when a new model release makes your carefully tuned deployment obsolete within months. It also breaks when quality slips and users quietly churn, a cost that never shows up on the GPU invoice.
The API case breaks at very high steady volume, under strict data rules and when rate limits cap growth. Keep your architecture able to move in either direction, and let measured usage, not intuition, trigger the switch.
Finally, remember opportunity cost. Engineer hours spent tuning inference servers are hours not spent on features customers pay for. For most early products, that trade only makes sense once cost is clearly threatening margins.
Frequently asked questions
- At what volume does self-hosting become cheaper than an API?
- There is no universal number. It depends on the model you need, the hardware, measured throughput, utilization, current API prices and your team’s time. Calculate it from your own logs and a real benchmark. Many teams find APIs cheaper until usage is both high and steady.
- Is running an LLM on my own laptop free?
- The marginal cost is small, just electricity and wear, which makes local models great for development and personal tools. It does not scale to serving customers, and smaller laptop-friendly models may not match the quality of larger hosted ones for demanding tasks.
- What hidden costs do teams miss when self-hosting?
- Idle GPU time, redundancy for uptime, monitoring, storage for model weights, network egress, security patching and engineer time for upgrades and incidents. Evaluation work when switching models is another. Together these often outweigh the savings on raw per-token price.
- How can I reduce API costs without self-hosting?
- Trim prompts, cache stable prefixes, route simple requests to smaller models, cap output length, batch non-urgent jobs and cache repeated answers. Measuring cost per feature, not just the monthly total, shows which part of the product is worth optimizing first.