← Resources
By Jordan Malik
— Staff Engineer, Inference
·
· GUIDE
When Self-Hosting LLMs Is Actually Cheaper Than an API
Self-hosting an LLM is cheaper than an API only above a token volume that depends entirely on GPU utilization — there's no single break-even number. This guide covers the real math, why batching and quantization decide it, the hidden operations cost, and why hybrid is usually the right answer.
There's no single break-even number
The honest answer to "when is self-hosting cheaper?" is a range gated by utilization, not a number — which is why published estimates disagree by orders of magnitude. They assume different GPU prices, utilization rates, and which API you're comparing against.
The useful shape: self-hosting tends to beat premium APIs at relatively modest sustained volume (single-digit millions of tokens a month on a well-used GPU), but beating cheap budget APIs requires far more (tens to hundreds of millions a month). A rough rule-of-thumb middle is a couple million tokens a day, sustained. Below that, an API almost always wins.
The GPU math
Self-hosting means renting or owning a GPU, and that fixed cost is what your token volume has to "fill." An on-demand H100 runs roughly $1.50–3.00/hour on specialized clouds (more on hyperscalers), so a dedicated 24/7 H100 is on the order of $1,000–5,000 a month depending on provider tier.
That monthly nut is fixed whether you use it or not — which is why the cost per million tokens is cheap only at high utilization. Well-utilized, a mid-size model can hit well under a dollar per million tokens; at 10% load, the real cost per token rises roughly tenfold, enough to make self-hosting more expensive than a premium API. An idle GPU is the fastest way to lose the cost argument.
Batching and quantization decide it
Two levers move the break-even more than anything else. Continuous batching (via modern inference servers) is the economic engine: serving many requests at once on the same GPU can cut cost per token dramatically — moving from one request at a time to large batches can improve throughput by orders of magnitude. Spiky, low-concurrency traffic underfills batches and runs several times worse than peak, which is why bursty workloads self-host poorly.
[Quantization](/articles/llm-quantization-explained-gguf) is the other lever: FP8 roughly doubles throughput at minimal quality loss, and INT4 can roughly triple it while shrinking the model enough to fit a smaller, cheaper GPU. Together, batching and quantization can swing cost per token several-fold.
The hidden cost: operations
The GPU bill is not the whole bill. Operations routinely add several times the raw hardware cost: an MLOps engineer is a six-figure salary, and each model update cycle — re-quantizing, testing, redeploying — is weeks of work. "Free" open-weight models can hide hundreds of thousands of dollars a year in engineering once you account for the inference server, driver and CUDA management, autoscaling, and monitoring.
This is the line item that sinks naive self-hosting math. The spec-sheet cost per token assumes the GPU runs itself; in reality, someone has to keep it running, and that person isn't free.
When the API clearly wins (and the non-cost factors)
An API is the right call for low or spiky volume, unpredictable traffic, teams without in-house MLOps, and any need for frontier-quality models you can't run yourself. By one estimate, an API is the better choice for the large majority of use cases — and API prices keep falling, which pushes the self-hosting break-even higher over time.
Two factors sit outside the cost math entirely: latency and privacy. Self-hosting wins for data residency, regulated workloads, and predictable in-region latency regardless of token economics — see our [on-premise guide](/articles/on-premise-ai-for-enterprise). Those can justify self-hosting even when the pure cost math doesn't.
Route by economics with osFoundry
The pragmatic answer is hybrid: route by economics, not ideology. High, steady volume belongs on local or owned inference, where filling a GPU's batches drives cost per token down. Mid-volume, predictable workloads fit a dedicated GPU endpoint — utilization-friendly throughput without owning hardware or carrying the full operations multiplier. Spiky, low, or unpredictable traffic, and anything needing frontier quality, stays on pay-per-token BYOK APIs.
osFoundry handles the levers that actually move the break-even — quantization, batching, and tier routing — automatically, and lets the same workload run local, on a dedicated endpoint, or through a BYOK API. So you capture the savings where self-hosting wins without managing CUDA, autoscaling, or duty-cycle math yourself. (Our [BYOK cost breakdown](/articles/byok-vs-managed-ai-cost-breakdown) covers the API side of the equation.)
Frequently asked questions
- At what token volume does self-hosting become cheaper than an API?
- It depends on utilization, so there's no single number. Roughly, self-hosting can beat premium APIs at single-digit millions of tokens a month on a well-used GPU, but beating cheap budget APIs takes tens to hundreds of millions a month. A common rule-of-thumb threshold is a couple million tokens a day, sustained — below that, an API usually wins.
- How much does it cost to run an H100 per month?
- An on-demand H100 is roughly $1.50–3.00/hour on specialized GPU clouds (higher on hyperscalers), so a dedicated 24/7 instance lands around $1,000–5,000 a month depending on the provider tier. That fixed cost is what your token volume must fill to make self-hosting pay off.
- Why does GPU utilization make or break self-hosting economics?
- Because the GPU costs the same whether it's busy or idle. Well-utilized, the cost per million tokens can drop below a dollar; at low load (say 10%), the real cost per token rises roughly tenfold, often making self-hosting more expensive than a premium API. An idle GPU is pure waste — utilization is the single biggest variable.
- Is self-hosting cheaper than a budget API like a hosted small model?
- Usually only at very high volume. Budget APIs are priced aggressively, so beating them by self-hosting typically requires tens to hundreds of millions of tokens a month at good utilization. Against premium APIs the crossover comes much sooner. Compare against the specific API you'd otherwise use, not a generic rate.
- How much does quantization reduce self-hosting cost?
- Substantially. FP8 roughly doubles throughput at minimal quality loss, and INT4 can roughly triple it while shrinking the model to fit a smaller, cheaper GPU. On a large model, that can cut cost per million tokens by half or more — quantization is one of the highest-leverage cost levers for self-hosting.
- What hidden costs come with self-hosting an LLM?
- Operations. Beyond the GPU bill, you pay for MLOps engineering (a six-figure salary), model-update cycles (weeks of work each), and the inference server, driver management, autoscaling, and monitoring. These routinely add several times the raw hardware cost and are what most naive break-even calculations leave out.
- Can I combine self-hosting and APIs to optimize cost?
- Yes — hybrid is usually the smartest approach. Run high, steady volume on local or owned inference, mid-volume predictable workloads on a dedicated endpoint, and spiky or low-volume traffic (plus anything needing frontier quality) on a pay-per-token API. Routing by economics captures the savings of self-hosting without paying for idle GPUs on bursty work.
Sources