TokenBill LLM cost statements

Prices updated 2026-10-05 09:30 KST

Cost guide

How Many Tokens Is 1,000 Words?

Updated 2026-10-05

The words-to-tokens conversion every LLM builder needs — plus the reverse table for reading context windows, and where the rule of thumb breaks.

The short answer

1,000 words of ordinary English is roughly 1,300 tokens. Going the other way, 1,000 tokens is roughly 750 words — about 4,000 characters. That is the whole answer for estimating.

The conversion comes from OpenAI’s documented rule of thumb: one token is about 4 characters, or about 0.75 English words. Flip it around and one word costs you about 1.33 tokens. Every number on this page is an approximation built on that ratio — fine for “will this fit” and “what will it roughly cost,” not fine for a hard limit or a bill you are going to rely on. We will get to exact counts below.

Words to tokens: the conversion table

For plain English prose, multiply your word count by ≈1.33. If you only have a character count, divide by 4.

You haveRoughly this many tokens
1 word1.3
100 words (a paragraph)130
500 words (about a page)650
1,000 words1,300
5,000 words (a long article)6,650
10,000 words (a short report)13,300
50,000 words (a novella)66,500
1,000 characters250
10,000 characters2,500

All figures approximate, for ordinary English prose with a typical subword tokenizer.

Tokens to words: how to read a context window

Model pages advertise context in tokens — 8K, 32K, 128K — and what you actually want to know is how much text that holds. Multiply tokens by ≈0.75 to get words, then divide by 500 for rough pages.

Token budgetRoughly this much English
4,096 tokens≈3,000 words · 6 pages
8,192 tokens≈6,000 words · 12 pages
32,768 tokens≈24,000 words · 50 pages
128,000 tokens≈96,000 words · 190 pages
1,000,000 tokens≈750,000 words · 1,500 pages

A “page” here means about 500 words of prose. Remember the budget covers everything: your prompt and the model’s reply share the same window.

When the 4-character rule breaks

Those tables assume ordinary prose. Tokenizers split text by statistical frequency, not by words, so anything unusual costs more tokens than the ratio suggests:

Code, JSON, and long URLs. Punctuation, braces, camelCase names, and random ID strings shatter into many small pieces. A dense JSON blob can easily cost noticeably more tokens per word than prose. If your workload is code-heavy, budget above the table — or better, measure with the real tokenizer.

Non-Latin scripts. This is the big one. In some tokenizers, Chinese, Japanese, Arabic, Hindi, and other non-Latin scripts can run several tokens per character, so the same meaning costs multiples of the English price. Multilingual apps should never budget on the English ratio.

Numbers and tables. Long columns of figures split in ways that look arbitrary — a spreadsheet pasted as text is not as cheap as its word count implies.

Different models, different tokenizers. A prompt is not a fixed number of tokens; it is a number per tokenizer. GPT-family models, Claude, and Gemini each count differently, so a token budget measured on one model does not transfer exactly to another. When precision matters, count with the tokenizer of the model you will actually call.

From words to dollars: a worked example

Tokens are the unit, but dollars are the question. The chain is: words × 1.33 → tokens, then (tokens ÷ 1,000,000) × price per 1M tokens.

Take a concrete case. You send a 1,000-word document for summarization: that is about 1,300 input tokens. At an illustrative price of $2.00 per million input tokens, one request costs 1,300 ÷ 1,000,000 × $2 ≈ $0.0026. Run that 100,000 times a month and you are looking at roughly $260/month — before the output tokens, which usually cost several times more than input. (This price is made up for the arithmetic; real rates move constantly.)

For real numbers, plug your word counts and today’s rates into the TokenBill calculator — it converts per-request tokens into monthly bills across 17 models, with caching and input/output mix included. Our cost guide covers the strategy side: why sticker prices mislead and which levers actually move the bill.

Prices move; verify current rates with the provider before deciding.

Getting an exact count

The rule of thumb answers “roughly how big.” When you need an exact number — a hard context limit, a cost estimate you will sign off on — count for real:

1. Use the provider’s own tokenizer tool or playground for the model you will call. 2. Check the API dashboard’s usage stats after a test run — most providers report exact input/output tokens per request. 3. For offline counting, use an open tokenizer library matching the model family (e.g. a tiktoken-compatible counter for GPT-family models).

One exact measurement of your real prompts beats a hundred table lookups. Log a sample of production requests, average them, and budget from that.

FAQ

Is a token the same as a word?
No. A token is a subword chunk — common words are often one token each, rare words split into pieces (the classic example: “unbelievable” becomes un · bel · iev · able, four tokens). In English prose the average works out to about 0.75 words per token.

How many tokens is one page of text?
About 500 words of prose is roughly 650 tokens. A 10-page document is on the order of 6,500 tokens.

Do different models count tokens differently?
Yes. Each model family uses its own tokenizer, so the same text produces different token counts on different models. Always count with the tokenizer of the model you plan to use.

How do I turn a token count into a cost?
Divide by one million and multiply by the model’s price per million tokens, separately for input and output. The TokenBill calculator does this across models with current prices.

Price your own text

Now that you can convert words to tokens, see what those tokens cost: run your per-request counts through the TokenBill calculator and compare monthly bills across models.