StringMash.com

Tokenizer

Every token in your text, and the numbers the model receives, with OpenAI's own tokenizers.

71 characters
Updates as you type
Tokens17
Characters71
Words9
Characters per token4.18
Spaces show as · and line breaks as ↵

Show the steps
  1. Counted with o200k_base, the tokenizer for GPT-4o, GPT-4.1, GPT-5, o3-mini and o4-mini.
  2. The first tokens: "Token" 4421, "ization" 2860, "·isn't" 12471, "·always" 3324, "·where" 1919, "·you'd" 35174, "·expect" 2665, ":" 25, "·" 220, "202" 1323, "6" 21, "," 11, ….
  3. A token can be a whole word, part of a word, a space plus a word, or punctuation. Common English words are usually one token each; rare words, numbers and other scripts take more.

Using the tokenizer

Type or paste text and it's split into tokens, separated by a bar, with spaces shown as · so you can see which token they belong to. Show: Token IDs gives the numbers instead, the actual input a model receives, and Show: The count gives just the total. Tokenizer switches between o200k, for GPT-4o, GPT-4.1 and GPT-5, and cl100k, for GPT-4 and GPT-3.5 Turbo.

Try the same sentence in both. o200k has twice the vocabulary, so it often needs fewer tokens for the same text.

What a token is

Language models don't read letters or words. They read tokens, chunks of text from a fixed vocabulary of about 100,000 entries in cl100k and 200,000 in o200k. A common word is usually one token, often with the space before it attached, so " world" is a single token. Rarer words split into pieces, numbers split into groups of up to three digits, and text in other scripts tends to take more tokens per character than English.

Tokens matter because they're what you pay for and what fills a model's context window. A prompt's cost and whether it fits are both measured in tokens, not characters.

How the splitting works

These tokenizers use byte pair encoding. Training starts from single bytes and repeatedly merges the most common adjacent pair into a new token, until the vocabulary is full. Encoding replays those merges on your text, in the order they were learned. Before that, a pattern cuts the text into pieces, so a merge never crosses from a word into punctuation or from letters into digits. That's why numbers come out in groups of up to three digits.

What about Claude, Gemini and Llama?

Each model family has its own tokenizer, and only some are published. The counts here are exact for OpenAI's models and only a rough guide for anyone else's. Anthropic offers a free token counting endpoint for Claude that returns its own estimate, and notes that Claude Opus 4.7 and later use a newer tokenizer that turns the same text into roughly 30 percent more tokens than earlier Claude models. That endpoint needs an API key and receives your text, so it isn't used here.

Your text stays here

The tokenizer runs in your browser. The first time you count, the page downloads the vocabulary for the tokenizer you've picked, a file of up to about 1 MB from this site, and it's cached after that. Nothing you type is sent anywhere.

Questions

What tokenizer does ChatGPT use?

The models behind it, such as GPT-4o and GPT-5, use o200k_base. Older GPT-4 and GPT-3.5 Turbo use cl100k_base.

Why is a space part of the next token?

The tokenizer learned that a word usually follows a space, so " the" with its space is one token. Showing spaces as · makes this visible.

Why do numbers take so many tokens?

The tokenizer splits digits into groups of at most three, so a long number becomes several tokens.

Sources

Added . What's new