Your country

Tools that support it use your country for local currency, number formats, units and paper size. Your choice is saved only in this browser.

Type a name or a two-letter code. Use the up and down arrow keys to move through the countries, Enter to choose one and Escape to close.

AI API Cost Calculator

Plan the monthly bill of an AI feature from requests, tokens, caching and batch.

Developer No upload Free, no sign-up

Your usage

10k and 1.5M work too.
System prompt, history, tools and the user’s text.
The visible answer.
Billed as output. OpenAI’s and Anthropic’s output_tokens already include them: count them once.
The part of the input read from the prompt cache.
Work that can wait for the Batch API.
Requests that create the cache instead of reading it.
Count tokens from a sample prompt and answer

Counting happens on this device; the first count downloads the tokenizer once.

Model and prices

    Prices per 1M tokens, US dollars

    An empty box means the provider lists no such price: cached tokens then cost the input price, and batch requests the standard price.

    Pricing page

    Save these prices as your own model

    Cost

    Cost per month —

    —Per request
    —Per day
    —Per month
    —Per year
    Cost a day by token type
    PartTokens a dayCost a day

      Enter a budget to see how many requests it pays for.
      Enter the rate your bank or card provider charges; none is built in.

      Compare models for this usage

      Providers to compare
      Models ranked by monthly cost for the usage above
      Model Per request Per month Per year vs cheapest

      Choose a model name to see its breakdown. Edited prices and your own models are compared too.

      Next steps

      Results are estimates for general information and planning, not financial advice. Banks and institutions may calculate differently (rounding, fees, rate changes). Confirm figures with your lender or a qualified adviser before deciding.

      About the AI API Cost Calculator

      Work out what an AI feature will cost before you build it. Enter how many requests it makes a day, how many input and output tokens each request has, how much of the input the prompt cache serves and how many requests can go through the Batch API, and pick a model: the calculator shows the cost per request, per day, per month and per year, and ranks the other models for the same workload.

      The presets hold the per-million-token prices that OpenAI, Anthropic, Google and Mistral AI publish for input, cached input, cache writes, output and batch requests, with the long-context prices that some models charge for very long prompts. Every price can be edited, and you can add your own model — a fine-tuned one, a negotiated rate or another provider. Paste a sample prompt and answer to count their tokens with the same engine as the LLM Token Counter.

      How to use it

      1. Enter the requests per day and the input and output tokens per request. Not sure about the tokens? Open Count tokens from a sample, paste a typical prompt and answer, and press Count and use.
      2. For reasoning models, add the average reasoning tokens per request: they are billed as output. If you copy the output tokens from an API’s usage figures, they already include the reasoning tokens, so count them only once.
      3. If your prompts repeat a long prefix (a system prompt, documents, tools), enter the cached share of the input. For a provider that charges for writing the cache, also enter how many cache writes happen a day.
      4. Enter the share of requests sent by batch for work that can wait for the Batch API.
      5. Choose a model. Its published prices fill the price boxes; change any of them for a discount or a new price, or save your own model.
      6. Read the cost per request, day, month and year, the breakdown and the notes. Enter a monthly budget to see how many requests it pays for, and compare every model in the table below — download it as CSV.

      Examples

      A support assistant on gpt-5.4-mini
      Input
      1,000 requests a day · 1,500 input and 400 output tokens · $0.75 input and $4.50 output per 1M tokens
      Result
      $0.00293 per request · $2.93 a day · $88.97 a month · $1,067.63 a year

      1,500 × $0.75 + 400 × $4.50 = $2,925 per million requests, or $0.002925 each.

      Caching a long prompt on Claude Sonnet 4.6
      Input
      1,000 requests a day · 10,000 input tokens, 80% from the cache · 500 output and 1,000 thinking tokens · 24 cache writes a day
      Result
      $31.56 a day (without caching: $52.50 a day)

      Cache reads cost $0.30 instead of $3 per million tokens; each 5-minute cache write costs $3.75 per million for the 8,000 cached tokens.

      Half the jobs through the Batch API on Claude Haiku 4.5
      Input
      200 requests a day · 2,000 input and 300 output tokens · 50% batch
      Result
      $0.00263 per request instead of $0.00350 · $0.525 a day

      Common uses

      • Checking that an AI feature pays for itself before you write the code, or before a pricing decision.
      • Choosing between a large and a small model, or between providers, for the same workload.
      • Seeing how much prompt caching and the Batch API save on a high-volume job.
      • Setting a monthly API budget and knowing how many requests it covers.

      How the cost is worked out

      Each token type has its own price per 1M tokens:

      • Cost per request = (input tokens not cached × input price + cached tokens × cached-input price + (output tokens + reasoning tokens) × output price) ÷ 1,000,000.
      • Requests sent through the Batch API use the batch prices; the share you enter is blended in.
      • A cache write replaces a cache read for the cached part of that request, at the cache-write price where the provider lists one and at the input price where it does not.
      • Per day = cost per request × requests per day + cache writes. Per year = per day × 365, and per month is a twelfth of the year (30.42 days).

      Prompt caching

      When many requests start with the same text — a long system prompt, a document, tool definitions — the provider can serve that part from its cache at a much lower price. The cached share is the part of the input tokens read from the cache. Anthropic charges extra to write the cache (1.25 times the input price for the 5-minute cache, 2 times for the 1-hour cache) and charges 0.1 times the input price to read it, less on some models, as its pricing page lists; some OpenAI models also list a cache-write price on the OpenAI pricing page. A request that finds no cache pays for writing it, so enter roughly how many times a day the cache is created.

      Batch requests and long prompts

      The Batch API returns results later instead of at once, at lower prices: Anthropic’s page states a 50% discount on input and output tokens, OpenAI and Google list batch prices on their pricing pages, and Mistral’s batch documentation says batch workloads run at a 50% discount. Some models charge more for very long prompts: OpenAI’s long-context prices apply to the whole request above 272,000 input tokens, and Gemini Pro models charge more above 200,000 tokens (Gemini API pricing). The calculator applies these automatically and says so in the notes.

      Reasoning and thinking tokens

      Reasoning models think before they answer, and those tokens cost money even though you never see them: OpenAI bills reasoning tokens as output tokens (reasoning guide), Google’s output prices include thinking tokens, and Anthropic reports thinking as part of the billed output tokens (extended thinking). Take the average from the usage figures of a few test calls: OpenAI reports output_tokens_details.reasoning_tokens and Anthropic output_tokens_details.thinking_tokens, and in both the output_tokens total already includes them. Enter the rest as output tokens and the reasoning separately, or the whole total as output with no reasoning tokens, but do not count them twice.

      Where the prices come from

      The presets are copies of the per-million-token prices on the providers’ own pages: OpenAI, Anthropic, Google Gemini API and Mistral AI, in US dollars. Prices that the provider lists as changing later, and sale prices, are marked, with a second preset at the later price or the list price. Prices change, so open the provider’s page before you commit to a budget — and edit any price the calculator shows.

      Limitations

      • The presets are copies of published list prices: check the provider’s page before you rely on them, and enter your own rate if you have a discount or credits.
      • Taxes, free tiers, regional or data-residency surcharges, priority or fast processing, tool fees (such as web search per call), images, audio and the per-hour storage of explicit context caches are not included.
      • Token counts of a sample are exact only for tokenizers that are published (OpenAI’s and the open models’); Claude and Gemini counts are the providers’ rule-of-thumb estimates.
      • Real requests also carry tokens you may not count: chat formatting, tool definitions, retrieved documents and retries. Compare with the usage report of a few test calls.
      • The calculator does not check a model’s context window: make sure your prompt and answer fit the model you choose.

      Privacy

      Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot. Counting a sample prompt downloads the tokenizer you pick once from MySmartCoPilot; your text and numbers never leave the browser.

      Frequently asked questions

      How do I know how many tokens my requests have?

      Paste a typical prompt and answer under Count tokens from a sample, or use the LLM Token Counter. After a few test calls, the usage figures the API returns give the exact numbers, including the tokens added by tools and chat formatting.

      Why would my real bill be higher than the estimate?

      Requests are often longer than the text you write: conversation history, tool definitions, retrieved documents and system instructions all count as input, reasoning models add hidden output tokens, and retries count twice. Taxes and surcharges are not included either. Check the usage figures of a few test calls and adjust the numbers.

      Which currency are the prices in?

      US dollars, as the providers bill. Enter the exchange rate from your bank or card provider under Show in your currency to see the amounts in your own currency as well.

      Can I use prices from another provider or my own discount?

      Yes. Edit any price box — the calculator marks the model as edited and can reset it — or save the prices as your own model, which stays in this browser and appears in the comparison.

      Is anything I enter sent anywhere?

      No. The calculation runs in your browser, and sample text is counted on your device. The only download is the tokenizer you choose, from this site, the first time you count a sample.

      Quick answers and tool search

      Type to search tools or to get a quick answer, for example 18% of 2500. Use the up and down arrow keys to move through the results, Enter to choose, and Escape to close.