Using the encoder
Type or paste text and you get its bytes in hex, ready to copy. The working lists every character with its code point and the bytes it became, and Separator chooses spaces, commas or nothing between bytes. Swap the boxes to decode bytes back into text.
How UTF-8 encodes a character
UTF-8 uses one to four bytes per character. Code points up to U+007F, the ASCII range, take one byte with the same value as in ASCII, which is why plain English text looks identical in ASCII and UTF-8. Up to U+07FF takes two bytes, up to U+FFFF three, and the rest, emoji included, four.
The first byte says how long the character is: 0xxxxxxx for one byte, 110xxxxx for two, 1110xxxx for three, 11110xxx for four. Every following byte starts 10, so a program can always tell where a character starts, even in the middle of a stream. é, U+00E9, becomes 11000011 10101001, which is C3 A9.
Bytes, not characters
A string's length in characters and its size in bytes are different numbers once anything beyond ASCII is involved. Café is four characters and five bytes, and a single emoji is four bytes. Databases, file formats and network protocols that limit size in bytes count the second number, and the encoder's working gives both.
Questions
How many bytes is a character in UTF-8?
One to four. ASCII characters take one, most accented Latin letters two, most other scripts three, and emoji four.
Is UTF-8 the same as Unicode?
No. Unicode numbers the characters; UTF-8 is one way to store those numbers as bytes. UTF-16 is another.
What's the difference between this and text to hex?
The bytes are the same. This page adds the per-character breakdown and a decoder that explains invalid bytes.
Sources
Added . What's new








