Using the decoder
Paste bytes in hex and the text appears. The decoder reads them however they were written: spaced or run together, with 0x in front, as \xE2 escapes, or as %E2 from a URL. The working shows which bytes made each character.
When the bytes aren't valid UTF-8
Not every sequence of bytes is valid UTF-8, and the decoder says which byte is wrong and why. A continuation byte (80 to BF) can't start a character. A lead byte must be followed by the right number of continuation bytes. A character can't be written with more bytes than it needs, a so-called overlong form, and C0, C1 and F5 to FF never appear at all. UTF-16 surrogates (U+D800 to U+DFFF) and anything past U+10FFFF are forbidden too.
By default the decoder stops at the first problem and explains it. Set Invalid bytes to Show as � to decode the rest anyway, with the replacement character in place of each bad byte, the way browsers display it.
Bytes that decode to the wrong thing
Sometimes the bytes are valid and the text still looks wrong, such as ’ where an apostrophe should be. That isn't a decoding error but text that was decoded with the wrong encoding somewhere earlier. The mojibake fixer repairs it.
Questions
What does � mean?
It's the replacement character, U+FFFD, shown where bytes couldn't be decoded. Once text contains it, the original bytes are gone.
Why is C0 AF invalid?
It's an overlong form of /, which is properly the single byte 2F. UTF-8 forbids overlong forms because they were used to slip characters past security checks.
Can I paste a URL-encoded string?
Yes. %E2%80%99 decodes to ’.
Sources
Added . What's new








