Using the token counter
Paste a prompt, a document or any text, and the count appears with the characters, the words and the characters per token. Tokenizer picks the vocabulary: o200k for GPT-4o, GPT-4.1, GPT-5 and the o3-mini and o4-mini models, cl100k for GPT-4, GPT-3.5 Turbo and the text-embedding-3 models. Show switches between the count, each token on its own, and the token IDs the model actually receives.
From also offers JSON to Token cost by format. Paste JSON and the same data is written as indented JSON, minified JSON, YAML, CSV and XML, each counted, so you can see which format fits the most data into a prompt.
What a token is
Language models don't read letters or words. They read tokens, chunks of text from a fixed vocabulary of about 100,000 entries in cl100k and 200,000 in o200k. A common word is usually one token, often with the space before it attached, so " world" is a single token. Rarer words split into pieces, numbers split into groups of up to three digits, and text in other scripts tends to take more tokens per character than English.
Tokens matter because they're what you pay for and what fills a model's context window. A prompt's cost and whether it fits are both measured in tokens, not characters.
Which format costs fewest tokens
For a list of records with the same fields, CSV is usually cheapest, because it writes each key once in the header instead of in every record. Minified JSON drops the indentation that indented JSON spends tokens on, and YAML drops most of the quotes and braces but still repeats every key. XML repeats each key twice, in the opening and closing tag. The comparison runs on your own data, since nesting, long strings and numbers all shift the balance.
What about Claude, Gemini and Llama?
Each model family has its own tokenizer, and only some are published. The counts here are exact for OpenAI's models and only a rough guide for anyone else's. Anthropic offers a free token counting endpoint for Claude that returns its own estimate, and notes that Claude Opus 4.7 and later use a newer tokenizer that turns the same text into roughly 30 percent more tokens than earlier Claude models. That endpoint needs an API key and receives your text, so it isn't used here.
Your text stays here
The tokenizer runs in your browser. The first time you count, the page downloads the vocabulary for the tokenizer you've picked, a file of up to about 1 MB from this site, and it's cached after that. Nothing you type is sent anywhere.
Questions
How many tokens is my text?
Paste it above. The count is exact for OpenAI models that use the tokenizer you've picked.
Which tokenizer does GPT-4o use?
o200k_base, as do GPT-4.1, GPT-5, o3-mini and o4-mini. GPT-4 and GPT-3.5 Turbo use cl100k_base.
Is a token a word?
Often, but not always. Common words are one token, rare words and numbers split into several, and a space usually joins the word after it.
Does this count tokens for Claude?
Not exactly. Claude's tokenizer isn't published, so use Anthropic's token counting endpoint for an exact figure. The counts here give a rough idea.
Sources
Added . What's new






