AI API cost calculator

API pricing is quoted per million tokens, which makes it impossible to feel. Enter what your app actually does per call and get a number in dollars — plus the cheapest model that could do the same work.

Prices verified 03 August 2026 against each provider’s official pricing page.

Where the money actually goes

Output tokens cost several times more than input tokens on every provider — commonly five times. That single fact reshapes how you build: stuffing a large prompt in is comparatively cheap, while asking for long prose back is what runs up a bill. If your costs are surprising you, the fix is almost always shortening responses, not shortening prompts.

The gap between tiers is bigger than people expect. The same workload can differ by fifty times between a frontier model and a small fast one, which is why routing matters: use the cheap model for classification, extraction, and formatting, and save the expensive one for the reasoning that actually needs it. The "cheapest listed model" row exists to show you that spread on your own numbers rather than in the abstract.

Two discounts worth knowing before you optimise anything else. Prompt caching charges roughly a tenth of the input rate for repeated context — if you send the same long system prompt or document every call, this is the single biggest lever available, and the cache row above shows its ceiling. Batch processing is typically half price on both directions when you can tolerate delayed results. Neither requires changing models.

Questions

How many tokens is my prompt?

Roughly one token per four characters of English, or about 0.75 words per token — so 1,000 words is near 1,300 tokens. Code and non-English text run higher. Every provider offers an exact tokenizer, but that estimate is close enough for budgeting.

Why is output so much more expensive than input?

Generating tokens is sequential and compute-bound, while reading your prompt is parallel. The economics follow the hardware. Practically: cap your response lengths, and ask for structured output rather than prose when you only need the data.

How current are these prices?

They were read from each provider's own pricing page on the date shown below the calculator. This is the fastest-moving data on this site, so it gets re-checked often — and note one scheduled change: Claude Sonnet 5's introductory rate rises on 1 September 2026.

Does this include caching and batch discounts?

The headline number is standard pricing. The cache-read row shows the floor if every input token were a cache hit, which is the realistic best case for apps with a large fixed prompt. Batch pricing is typically half of standard on both input and output.

Related calculators