Skip to content
mathbehind

Self-Host vs API LLM Cost Calculator

Compare renting GPUs to run an open model with paying per token for an API model, and find the monthly volume where self-hosting starts to pay.

Your numbers

Prices and fees in this tool are in US dollars.

M tokens
M tokens
$

From your cloud or GPU provider for the GPU you need. The Local LLM VRAM Calculator shows how much GPU memory a model needs.

hrs

730 is always on. Fewer if you shut down outside working hours.

Measure this with your model, server and batch size. It varies widely.

$

Storage, networking, monitoring and a share of engineering time.

Cheaper option at this volume

API

Monthly difference

$1,245

API cost per month
$880.00
Self-hosting cost per month
$2,125.00
Break-even volume (at your input/output mix)
5,795 M tokens
GPU capacity used
36.5%
Capacity check
Within capacity
The math behind it
  1. API

    2,000M × $0.20 + 400M × $1.20equals$880.00

  2. Self-hosting

    $2.50 × 1 GPU × 730 h + $300equals$2,125.00

  3. Difference

    |$880.00 − $2,125.00|equals$1,245

  4. Break-even

    $2,125.00 ÷ $0.367 per Mequals5,795M tokens

Running an open model on rented GPUs has a fixed monthly cost, while an API charges for every token. At low volume the API is almost always cheaper; at high, steady volume self-hosting can win. This calculator compares the two for your workload, shows the monthly token volume where they cost the same, and checks whether your GPUs can actually handle the traffic.

Figures checked by Muhammad Ahmad against OpenAI API pricing, Claude Platform Docs: Pricing and Gemini API pricing. Prices and rules change, so confirm with the official source before relying on them. How we check the math

How it works

API cost = input tokens (in millions) × the model's input price + output tokens (in millions) × its output price, using the verified prices from the AI API Cost Calculator.

Self-hosting cost = GPU price per hour × number of GPUs × hours running + other monthly costs such as storage, networking, monitoring and engineering time. It stays the same whether the GPUs are busy or idle.

Break-even volume = self-hosting cost ÷ the API's average price per million tokens at your input/output mix. Below it the API is cheaper; above it self-hosting is, as long as the GPUs have capacity.

Capacity = tokens per second per GPU × GPUs × hours × 3,600. If your monthly tokens exceed it, you need more GPUs, which raises the self-hosting cost. Real throughput depends on the model, server software, batch size and prompt lengths, so measure it rather than guessing.

A worked example

2,000 million input and 400 million output tokens a month on GPT-5.6 Luna ($0.20 in, $1.20 out) cost $880 through the API. One GPU at $2.50 an hour running all month costs $1,825, plus $300 of other costs: $2,125. The API is $1,245 a month cheaper. Its average price is about $0.37 per million tokens, so self-hosting only breaks even at about 5,795 million tokens a month. This GPU can process about 6,570 million a month at 2,500 tokens a second, so breaking even would keep it about 88% busy around the clock, with little room for peaks.

Questions people ask

When is self-hosting an LLM cheaper than an API?

When your volume is high and steady enough to keep the GPUs busy, and the API model you'd otherwise use is expensive. Against cheap API models, the break-even volume is often very high.

What costs does self-hosting add besides GPUs?

Storage for model weights, networking, monitoring, scaling for peaks, security updates and engineer time. These don't appear on an API bill, so include a realistic figure in other monthly costs.

How many tokens per second can one GPU handle?

It varies widely with model size, precision, server software and how many requests are batched together. Benchmark your own model and settings; published figures are rarely comparable.

Is the quality the same?

Not necessarily. An open model that fits on one GPU may be weaker than the API model you're comparing with. Compare quality on your own tasks before comparing cost.

Can I turn GPUs off to save money?

Yes, if traffic allows. Running only during working hours cuts the GPU bill sharply, but requests outside those hours need a fallback, often the API itself.