AI API Cost Calculator
Plan the monthly bill of an AI feature from requests, tokens, caching and batch.
Your usage
Count tokens from a sample prompt and answer
Counting happens on this device; the first count downloads the tokenizer once.
Model and prices
Edited pricesPrices per 1M tokens, US dollars
An empty box means the provider lists no such price: cached tokens then cost the input price, and batch requests the standard price.
Save these prices as your own model
Cost
| Part | Tokens a day | Cost a day |
|---|
Compare models for this usage
| Model | Per request | Per month | Per year | vs cheapest |
|---|
Choose a model name to see its breakdown. Edited prices and your own models are compared too.
Results are estimates for general information and planning, not financial advice. Banks and institutions may calculate differently (rounding, fees, rate changes). Confirm figures with your lender or a qualified adviser before deciding.
About the AI API Cost Calculator
Work out what an AI feature will cost before you build it. Enter how many requests it makes a day, how many input and output tokens each request has, how much of the input the prompt cache serves and how many requests can go through the Batch API, and pick a model: the calculator shows the cost per request, per day, per month and per year, and ranks the other models for the same workload.
The presets hold the per-million-token prices that OpenAI, Anthropic, Google and Mistral AI publish for input, cached input, cache writes, output and batch requests, with the long-context prices that some models charge for very long prompts. Every price can be edited, and you can add your own model — a fine-tuned one, a negotiated rate or another provider. Paste a sample prompt and answer to count their tokens with the same engine as the LLM Token Counter.
How to use it
- Enter the requests per day and the input and output tokens per request. Not sure about the tokens? Open Count tokens from a sample, paste a typical prompt and answer, and press Count and use.
- For reasoning models, add the average reasoning tokens per request: they are billed as output. If you copy the output tokens from an API’s usage figures, they already include the reasoning tokens, so count them only once.
- If your prompts repeat a long prefix (a system prompt, documents, tools), enter the cached share of the input. For a provider that charges for writing the cache, also enter how many cache writes happen a day.
- Enter the share of requests sent by batch for work that can wait for the Batch API.
- Choose a model. Its published prices fill the price boxes; change any of them for a discount or a new price, or save your own model.
- Read the cost per request, day, month and year, the breakdown and the notes. Enter a monthly budget to see how many requests it pays for, and compare every model in the table below — download it as CSV.
Examples
1,000 requests a day · 1,500 input and 400 output tokens · $0.75 input and $4.50 output per 1M tokens
$0.00293 per request · $2.93 a day · $88.97 a month · $1,067.63 a year
1,500 × $0.75 + 400 × $4.50 = $2,925 per million requests, or $0.002925 each.
1,000 requests a day · 10,000 input tokens, 80% from the cache · 500 output and 1,000 thinking tokens · 24 cache writes a day
$31.56 a day (without caching: $52.50 a day)
Cache reads cost $0.30 instead of $3 per million tokens; each 5-minute cache write costs $3.75 per million for the 8,000 cached tokens.
200 requests a day · 2,000 input and 300 output tokens · 50% batch
$0.00263 per request instead of $0.00350 · $0.525 a day
Common uses
- Checking that an AI feature pays for itself before you write the code, or before a pricing decision.
- Choosing between a large and a small model, or between providers, for the same workload.
- Seeing how much prompt caching and the Batch API save on a high-volume job.
- Setting a monthly API budget and knowing how many requests it covers.
How the cost is worked out
Each token type has its own price per 1M tokens:
- Cost per request = (input tokens not cached × input price + cached tokens × cached-input price + (output tokens + reasoning tokens) × output price) ÷ 1,000,000.
- Requests sent through the Batch API use the batch prices; the share you enter is blended in.
- A cache write replaces a cache read for the cached part of that request, at the cache-write price where the provider lists one and at the input price where it does not.
- Per day = cost per request × requests per day + cache writes. Per year = per day × 365, and per month is a twelfth of the year (30.42 days).
Prompt caching
When many requests start with the same text — a long system prompt, a document, tool definitions — the provider can serve that part from its cache at a much lower price. The cached share is the part of the input tokens read from the cache. Anthropic charges extra to write the cache (1.25 times the input price for the 5-minute cache, 2 times for the 1-hour cache) and charges 0.1 times the input price to read it, less on some models, as its pricing page lists; some OpenAI models also list a cache-write price on the OpenAI pricing page. A request that finds no cache pays for writing it, so enter roughly how many times a day the cache is created.
Batch requests and long prompts
The Batch API returns results later instead of at once, at lower prices: Anthropic’s page states a 50% discount on input and output tokens, OpenAI and Google list batch prices on their pricing pages, and Mistral’s batch documentation says batch workloads run at a 50% discount. Some models charge more for very long prompts: OpenAI’s long-context prices apply to the whole request above 272,000 input tokens, and Gemini Pro models charge more above 200,000 tokens (Gemini API pricing). The calculator applies these automatically and says so in the notes.
Reasoning and thinking tokens
Reasoning models think before they answer, and those tokens cost money even though you never see them: OpenAI bills reasoning tokens as output tokens (reasoning guide), Google’s output prices include thinking tokens, and Anthropic reports thinking as part of the billed output tokens (extended thinking). Take the average from the usage figures of a few test calls: OpenAI reports output_tokens_details.reasoning_tokens and Anthropic output_tokens_details.thinking_tokens, and in both the output_tokens total already includes them. Enter the rest as output tokens and the reasoning separately, or the whole total as output with no reasoning tokens, but do not count them twice.
Where the prices come from
The presets are copies of the per-million-token prices on the providers’ own pages: OpenAI, Anthropic, Google Gemini API and Mistral AI, in US dollars. Prices that the provider lists as changing later, and sale prices, are marked, with a second preset at the later price or the list price. Prices change, so open the provider’s page before you commit to a budget — and edit any price the calculator shows.
Limitations
- The presets are copies of published list prices: check the provider’s page before you rely on them, and enter your own rate if you have a discount or credits.
- Taxes, free tiers, regional or data-residency surcharges, priority or fast processing, tool fees (such as web search per call), images, audio and the per-hour storage of explicit context caches are not included.
- Token counts of a sample are exact only for tokenizers that are published (OpenAI’s and the open models’); Claude and Gemini counts are the providers’ rule-of-thumb estimates.
- Real requests also carry tokens you may not count: chat formatting, tool definitions, retrieved documents and retries. Compare with the usage report of a few test calls.
- The calculator does not check a model’s context window: make sure your prompt and answer fit the model you choose.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot. Counting a sample prompt downloads the tokenizer you pick once from MySmartCoPilot; your text and numbers never leave the browser.
Frequently asked questions
How do I know how many tokens my requests have?
Paste a typical prompt and answer under Count tokens from a sample, or use the LLM Token Counter. After a few test calls, the usage figures the API returns give the exact numbers, including the tokens added by tools and chat formatting.
Why would my real bill be higher than the estimate?
Requests are often longer than the text you write: conversation history, tool definitions, retrieved documents and system instructions all count as input, reasoning models add hidden output tokens, and retries count twice. Taxes and surcharges are not included either. Check the usage figures of a few test calls and adjust the numbers.
Which currency are the prices in?
US dollars, as the providers bill. Enter the exchange rate from your bank or card provider under Show in your currency to see the amounts in your own currency as well.
Can I use prices from another provider or my own discount?
Yes. Edit any price box — the calculator marks the model as edited and can reset it — or save the prices as your own model, which stays in this browser and appears in the comparison.
Is anything I enter sent anywhere?
No. The calculation runs in your browser, and sample text is counted on your device. The only download is the tokenizer you choose, from this site, the first time you count a sample.