TokenBill LLM price desk

Prices updated 2026-10-03 12:01 KST

LLM API Cost Guide

How to estimate what a model really costs — and pick the cheapest one for your workload.

1. Token prices lie

Most people compare the "$/1M tokens" sticker price. That number alone tells you almost nothing about your bill. A support chatbot and a coding agent have completely different shapes: one sends short prompts and gets short replies millions of times, the other sends 50,000-token codebases a few hundred times. The same model can be the cheapest option for one workload and the most expensive for another. Always compute cost per task × tasks per month, not cost per token. That's what the TokenBill calculator does.

2. Input vs. output asymmetry

Output tokens typically cost 3–5× more than input tokens. Workloads that generate long responses (drafting, summarization, code generation) are output-heavy and punish expensive output pricing. Workloads like classification or RAG search are input-heavy and care far more about input price. Before choosing a model, measure your real input/output ratio from logs — most teams guess wrong.

3. Caching is the biggest lever

Prompt caching (reusing the prefix of a prompt across requests) can cut input costs by 50–90% depending on the provider. If your app sends the same system prompt, few-shot examples, or retrieved documents repeatedly, your cache hit rate is the single most important number in your cost model. A "cheap" model with no caching support can easily lose to a pricier model with a 70% cache hit rate. Enter your real hit rate in the calculator — the default 50% is only a starting guess.

Worked example: pricing a support bot

Take a concrete workload: 10,000 support replies per month, each averaging 500 input tokens and 300 output tokens, with a 50% cache hit rate and the standard 90% cached-input discount. Using the prices listed on TokenBill on 2026-10-03:

GPT-5.6 Luna — $4.15/mo  ·  Claude Sonnet 5.5 — $35.50/mo  ·  GPT-6 Astra — $177.50/mo

The same replies cost 43× more on the most expensive model than the cheapest — before any quality comparison. The math per request is simple: (500 × input price × 0.55 + 300 × output price) ÷ 1,000,000, where 0.55 is the cache adjustment (1 − 0.5 × 0.9). Run your own numbers through the calculator — the ranking flips as soon as the workload shape changes.

Prices move; verify current rates with the provider before deciding.

4. Don't over-buy intelligence

Frontier models cost an order of magnitude more than mid-tier ones. For many production tasks — classification, extraction, routing, first-draft support replies — a smaller model scores nearly as well at a fraction of the price. The practical pattern is tiering: use the cheap model by default, escalate only the hard cases. Measure quality on your own evals, not on benchmark leaderboards.

5. Re-check prices quarterly

LLM pricing moves fast: providers cut rates, add cached-input tiers, and launch cheaper model generations every few months. A model choice that was optimal in January can be 2× overpriced by June. Put a recurring calendar reminder to re-run your workloads through the calculator with fresh prices.

FAQ

How much does it cost to run a chatbot on an LLM API per month?
It depends on request volume, tokens per request, and model choice — the worked example above ranges from about $4 to about $178 per month for the same 10,000 replies. Plug your own numbers into the calculator for a real answer.

Why is my LLM bill higher than the token price suggested?
Three usual culprits: output tokens cost 3–5× input tokens; scaffolding (system prompts, RAG context, conversation history) inflates input far beyond the user message; and retries, tool calls, and reasoning tokens each add paid requests.

Do cached input tokens really cost less?
Yes. Providers discount repeated prompt prefixes by roughly 50–90%. If your app re-sends the same system prompt or documents, cache hit rate is the single biggest lever on your bill.

How often should I re-check model prices?
Providers cut rates and launch cheaper model generations every few months. Re-run your workloads quarterly — TokenBill refreshes its price table regularly so the comparison stays current.

Is a cheaper model always worse?
No. For classification, extraction, routing, and first-draft replies, smaller models often match frontier quality at a fraction of the price. The practical pattern is tiering: cheap model by default, escalate only hard cases — and validate on your own evals, not leaderboards.

Quick checklist

① Measure your requests/month, average input tokens, and average output tokens from real logs. ② Estimate your cache hit rate (or measure it — most providers show it in the dashboard). ③ Compare 3–5 candidate models on your numbers, not sticker prices. ④ Validate quality on your own task before switching. ⑤ Re-check every quarter.