Using the tokenizer
Type or paste text and it's split into tokens, separated by a bar, with spaces shown as · so you can see which token they belong to. Show: Token IDs gives the numbers instead, the actual input a model receives, and Show: The count gives just the total. Tokenizer switches between o200k, for GPT-4o, GPT-4.1 and GPT-5, and cl100k, for GPT-4 and GPT-3.5 Turbo.
Try the same sentence in both. o200k has twice the vocabulary, so it often needs fewer tokens for the same text.
What a token is
Language models don't read letters or words. They read tokens, chunks of text from a fixed vocabulary of about 100,000 entries in cl100k and 200,000 in o200k. A common word is usually one token, often with the space before it attached, so " world" is a single token. Rarer words split into pieces, numbers split into groups of up to three digits, and text in other scripts tends to take more tokens per character than English.
Tokens matter because they're what you pay for and what fills a model's context window. A prompt's cost and whether it fits are both measured in tokens, not characters.
How the splitting works
These tokenizers use byte pair encoding. Training starts from single bytes and repeatedly merges the most common adjacent pair into a new token, until the vocabulary is full. Encoding replays those merges on your text, in the order they were learned. Before that, a pattern cuts the text into pieces, so a merge never crosses from a word into punctuation or from letters into digits. That's why numbers come out in groups of up to three digits.
What about Claude, Gemini and Llama?
Each model family has its own tokenizer, and only some are published. The counts here are exact for OpenAI's models and only a rough guide for anyone else's. Anthropic offers a free token counting endpoint for Claude that returns its own estimate, and notes that Claude Opus 4.7 and later use a newer tokenizer that turns the same text into roughly 30 percent more tokens than earlier Claude models. That endpoint needs an API key and receives your text, so it isn't used here.
Your text stays here
The tokenizer runs in your browser. The first time you count, the page downloads the vocabulary for the tokenizer you've picked, a file of up to about 1 MB from this site, and it's cached after that. Nothing you type is sent anywhere.
Questions
What tokenizer does ChatGPT use?
The models behind it, such as GPT-4o and GPT-5, use o200k_base. Older GPT-4 and GPT-3.5 Turbo use cl100k_base.
Why is a space part of the next token?
The tokenizer learned that a word usually follows a space, so " the" with its space is one token. Showing spaces as · makes this visible.
Why do numbers take so many tokens?
The tokenizer splits digits into groups of at most three, so a long number becomes several tokens.
Sources
Added . What's new






