LLM Cost Calculator: API vs Self-Hosting
At what point does running your own GPU beat paying per token? Enter your volume and prices — get the monthly cost each way and the break-even point. You supply the prices, so the math reflects your real rates.
Pick how you actually pay for the GPU. Every rate below is yours to edit — nothing is fetched or assumed.
Default is the ECB reference rate of 18 Aug 2026. Set the rate you are actually billed at — this tool never looks FX up, and vendors who publish their own USD price rarely use the market rate.
One all-in figure, if you already know it.
Setup is billed once, so month one costs more than the average this shows. If you might cancel inside a year, set the spread to how long you will actually keep the box.
Resale value at the end of the amortization window is not modeled, so this reads slightly high for hardware you plan to sell on.
Fills the fields above so you can adjust from a real starting point. Rates verified 19 Aug 2026 from each vendor's own pricing data — they change, so check before you commit.
Two things these buttons deliberately do not do. They do not name a cheapest provider — the tiers are not equivalent, and a "community" pod is not the same product as a dedicated box. And they do not pick a GPU for your model: check the model fits first with the VRAM calculator. Hetzner's figures are list prices before VAT, which is applied by billing country, so a German buyer pays more than the number shown.
Related guides
- Self-hosting LLMs vs cloud APIs — the full cost, performance, and privacy comparison
- Production LLM cost optimization — cut the token bill before you switch
- LLM VRAM calculator — check the model fits the GPU you're pricing
Related reading
Self-Hosting LLMs vs Cloud APIs: Cost, Speed, Privacy 2026
Self-hosting LLMs with Ollama, vLLM, and llama.cpp vs cloud APIs: vendor-cited rates, hardware and power costs, payback periods, and when each wins.
Read →Hetzner Cloud for AI: GPU Server Setup and Cost Guide 2026
Run AI on Hetzner Cloud: €5.49/mo CPU instances to €184/mo RTX 4000 Ada GPU servers. Post-June-2026 pricing, setup, and a sourced AWS/GCP comparison.
Read →Semantic Caching vs. Prompt Caching: Where Break-Even Sits
Two caching layers on one workload, priced from Anthropic's official cache rates. Where each pays, and the measurement that tells you which fits.
Read →Find these tools useful?
Get one short weekly dispatch with new tools, guides, and what we tested. No spam, unsubscribe anytime.