StringMash.com

Token counter

Exact counts for OpenAI's models, worked out on your device, and the cheapest format for your data.

Conversion

A count can't be turned back into text. Choose Show: token IDs to see the numbers the text becomes.

59 characters
Updates as you type
Tokens13
Characters59
Words11
Characters per token4.54
Counted in your browser

Show the steps
  1. Counted with o200k_base, the tokenizer for GPT-4o, GPT-4.1, GPT-5, o3-mini and o4-mini.
  2. The first tokens: "Tokens" 30325, "·are" 553, "·what" 1412, "·language" 6439, "·models" 7015, "·read" 1729, "," 11, "·and" 326, "·what" 1412, "·you" 481, "·pay" 2777, "·for" 395, ….
  3. A token can be a whole word, part of a word, a space plus a word, or punctuation. Common English words are usually one token each; rare words, numbers and other scripts take more.

Using the token counter

Paste a prompt, a document or any text, and the count appears with the characters, the words and the characters per token. Tokenizer picks the vocabulary: o200k for GPT-4o, GPT-4.1, GPT-5 and the o3-mini and o4-mini models, cl100k for GPT-4, GPT-3.5 Turbo and the text-embedding-3 models. Show switches between the count, each token on its own, and the token IDs the model actually receives.

From also offers JSON to Token cost by format. Paste JSON and the same data is written as indented JSON, minified JSON, YAML, CSV and XML, each counted, so you can see which format fits the most data into a prompt.

What a token is

Language models don't read letters or words. They read tokens, chunks of text from a fixed vocabulary of about 100,000 entries in cl100k and 200,000 in o200k. A common word is usually one token, often with the space before it attached, so " world" is a single token. Rarer words split into pieces, numbers split into groups of up to three digits, and text in other scripts tends to take more tokens per character than English.

Tokens matter because they're what you pay for and what fills a model's context window. A prompt's cost and whether it fits are both measured in tokens, not characters.

Which format costs fewest tokens

For a list of records with the same fields, CSV is usually cheapest, because it writes each key once in the header instead of in every record. Minified JSON drops the indentation that indented JSON spends tokens on, and YAML drops most of the quotes and braces but still repeats every key. XML repeats each key twice, in the opening and closing tag. The comparison runs on your own data, since nesting, long strings and numbers all shift the balance.

What about Claude, Gemini and Llama?

Each model family has its own tokenizer, and only some are published. The counts here are exact for OpenAI's models and only a rough guide for anyone else's. Anthropic offers a free token counting endpoint for Claude that returns its own estimate, and notes that Claude Opus 4.7 and later use a newer tokenizer that turns the same text into roughly 30 percent more tokens than earlier Claude models. That endpoint needs an API key and receives your text, so it isn't used here.

Your text stays here

The tokenizer runs in your browser. The first time you count, the page downloads the vocabulary for the tokenizer you've picked, a file of up to about 1 MB from this site, and it's cached after that. Nothing you type is sent anywhere.

Questions

How many tokens is my text?

Paste it above. The count is exact for OpenAI models that use the tokenizer you've picked.

Which tokenizer does GPT-4o use?

o200k_base, as do GPT-4.1, GPT-5, o3-mini and o4-mini. GPT-4 and GPT-3.5 Turbo use cl100k_base.

Is a token a word?

Often, but not always. Common words are one token, rare words and numbers split into several, and a space usually joins the word after it.

Does this count tokens for Claude?

Not exactly. Claude's tokenizer isn't published, so use Anthropic's token counting endpoint for an exact figure. The counts here give a rough idea.

Sources

Added . What's new