LLM Token Counter
Exact token counts for OpenAI and open models — your text never leaves the browser.
Or drop a text file on the box. Counting happens on this device; nothing is uploaded.
Context window
Use your model’s documented limit; 128k and 1M work too.
—
Cost estimate
The output price applies to the tokens reserved for the reply. Prices come from your provider’s price list.
Enter a price to see the cost.
Tokens
Tokens appear here in alternating colours. ‹E0 A4› marks a token that holds only part of a character.
Compare tokenizers
Counts the same text with every tokenizer. Tokenizers you have not used yet are downloaded first (up to 11.2 MB in total).
| Tokenizer | Tokens | vs selected |
|---|
Open-model tokenizer files: Qwen and Mistral under the Apache 2.0 licence, DeepSeek under the MIT licence.
About the LLM Token Counter
Language models read text as tokens — words, parts of words, punctuation and bytes — and context limits and API prices are counted in them. Paste a prompt, a document or code to get the exact number of tokens for OpenAI’s encodings (o200k_base for GPT-4o, GPT-4.1, GPT-5 and the o-series; cl100k_base for GPT-4 and GPT-3.5; o200k_harmony for gpt-oss) and for open models whose tokenizers are published: Qwen, Mistral and DeepSeek.
Every token is shown in colour, so you can see where the text is split. The page also shows how much of a context window the text fills and, from a price you enter, what it costs to send. Anthropic and Google do not publish the tokenizers of Claude and Gemini, so for those models the page gives a clearly labelled estimate based on the providers’ own rules of thumb.
How to use it
- Paste or type your text, open a text file, or click Load sample.
- Choose the tokenizer of your model. The list says which models use each one; the first use downloads it once.
- Read the token count, then look at the coloured tokens below (or switch to IDs to see and copy the token numbers).
- Enter your model’s context window and, if you like, how many tokens to keep free for the reply, to see whether the text fits.
- Enter the price per million tokens from your provider’s price list to see the cost per request and for many requests.
- To see how much the choice of model matters, click Compare all to count the same text with every tokenizer.
Examples
Hello world
o200k_base: 2 tokens (13225, 2375) cl100k_base: 2 tokens (9906, 1917) Mistral v3: 2 tokens (23325, 2294)
Token numbers differ between tokenizers even when the count is the same.
नमस्ते दुनिया
o200k_base: 5 tokens · cl100k_base: 13 tokens · Qwen3.5: 6 tokens
Newer tokenizers with larger vocabularies need far fewer tokens for Indian languages.
12,000 tokens at $2.50 per million input tokens, 500 requests
$0.03 per request · $15.00 for 500 requests
Common uses
- Checking that a long prompt, document or chat history fits a model’s context window before sending it.
- Estimating what an API job costs before running it on thousands of requests.
- Comparing how many tokens the same text needs in different models — especially for Hindi and other non-English text.
- Seeing exactly how a tokenizer splits words, numbers, code and emoji while writing prompts.
Which tokenizer does my model use?
- o200k_base: GPT-5 models, GPT-4.5, GPT-4.1, GPT-4o, o1, o3 and o4-mini.
- cl100k_base: GPT-4, GPT-3.5 Turbo and the text-embedding-3 and ada-002 embedding models.
- o200k_harmony: gpt-oss-120b and gpt-oss-20b (the same tokens as o200k_base, plus the harmony chat-format tokens).
- Qwen3.5 / 3.6 / 3.8 and Qwen3 / Qwen2.5: two different vocabularies (about 248,000 and 151,000 tokens).
- Mistral Tekken: Mistral Small 4, Medium 3.5, Large 3, Ministral 3, Devstral 2 and Mistral Nemo. Mistral v3: Mistral 7B v0.3 and Mixtral 8x22B.
- DeepSeek-V3 / R1.
The OpenAI list follows OpenAI’s own tiktoken library. For the open models, each family was checked: the models named give the same tokens for ordinary text.
Claude and Gemini: estimates only
Anthropic and Google do not publish the tokenizers of Claude and Gemini, so no website can count their tokens exactly offline. The estimates here use the providers’ own rules of thumb: Anthropic says a Claude token is about 3.5 English characters, and that the newer tokenizer introduced with Claude Opus 4.7 gives roughly 30% more tokens for the same text (1 to 1.35 times as many, depending on the content); Google says a Gemini token is about 4 characters. Real counts differ, most of all for code and non-English text.
Choose Claude, newer tokenizer for Claude Opus 4.7, 4.8, 5 and 5.5, Sonnet 5 and 5.5, Fable 5 and 5.1 and Mythos 5 and 5.1, and Claude, older tokenizer for Opus 4.6, Sonnet 4.6, Haiku 4.5 and older models. For an exact number, use the providers’ token-counting API endpoints (they need an API key; Anthropic’s is free of charge).
What the count includes
The count is the tokens of your text alone. Chat APIs add a few tokens per message for roles and separators, and more for tool definitions, images and files, so the number on your bill can be a little higher. Some open models add a start token (like <s>) that is not counted here.
Special tokens such as <|endoftext|>: OpenAI encodings count them as ordinary text unless you tick Read special tokens as single tokens (the way the model’s own chat format writes them). The Qwen, Mistral and DeepSeek tokenizers always treat their own special tokens as single tokens, as their reference tokenizers do.
How exact is it?
The OpenAI encodings run OpenAI’s published tiktoken vocabularies (through the open-source gpt-tokenizer library); the Qwen, Mistral and DeepSeek counts run each model’s own tokenizer.json file. The token numbers were checked against OpenAI’s tiktoken and Hugging Face’s reference tokenizers on English, code, Hindi, Chinese, Japanese, Korean, Arabic, emoji and a 200,000-character book extract, with identical results.
Context windows and cost
A context window is the most tokens a model can handle in one request — your text plus the reply. Enter the reply length under Reserve for the reply to see what is left. Write large numbers in full or with k and M (128k = 128,000; 1M = 1,000,000). Check your provider’s current documentation for the window and prices: they change, and many providers charge different prices for input, output and cached tokens.
Where the tokenizers come from
OpenAI’s encodings come from the gpt-tokenizer package (MIT licence). The Qwen tokenizer files are © Alibaba Cloud and the Mistral files © Mistral AI, both under the Apache 2.0 licence; the DeepSeek file is © DeepSeek under the MIT licence. They are served from this site, unchanged except for removed whitespace and compression; the licence texts and a notice listing the source of each file are linked under Compare tokenizers.
Limitations
- Claude and Gemini counts are estimates (see above), not exact.
- Token boundaries are shown for the first 5,000 tokens; the count and the copied IDs always cover the whole text.
- Very long texts (several MB) take a few seconds to count, and each open-model tokenizer is a 0.4–3.3 MB download the first time.
- Images, audio, files and tool definitions are counted by the providers in their own ways; only text is counted here.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot. The tokenizer you choose is downloaded once from MySmartCoPilot (0.4–3.3 MB); your text is never sent anywhere.
Frequently asked questions
How many words is 1,000 tokens?
About 750 English words: OpenAI’s rule of thumb is that a token is about ¾ of a word, and Google gives 100 tokens ≈ 60–80 English words for Gemini. Other languages and code use more tokens per word — paste your own text to see the real number.
Does the count match my OpenAI bill?
For the text itself, yes — the same encodings are used. The API also adds a few tokens per chat message and counts tool definitions and images, so billed input tokens are usually a little higher than the count of your text.
Why do Hindi and other Indian languages use so many tokens?
Tokenizers learn their vocabulary mostly from English text, so other scripts are split into smaller pieces, sometimes single bytes. Newer tokenizers with larger vocabularies (o200k_base, Qwen3.5) need far fewer tokens for Hindi than older ones like cl100k_base. Paste your text and use Compare all to see the difference.
What does a token like ‹E0 A4› mean?
Some tokens hold only part of a character: a Devanagari letter or an emoji takes 3–4 bytes in UTF-8, and the tokenizer may split those bytes. Such a token is shown by its bytes; the character appears in the token that completes it.
Is my text uploaded?
No. Counting runs in your browser. The only download is the tokenizer itself, from this site, the first time you choose it.