The short answer
ASCII is a set of 128 characters, the English letters, digits and punctuation, each stored in one byte. Unicode is a catalogue that gives a number, a code point, to every character in every script, more than 150,000 of them. UTF-8 and UTF-16 are encodings: ways of storing those code points as bytes. UTF-8 uses one to four bytes per character and is identical to ASCII for ASCII text; UTF-16 uses two or four.
Look inside any character
Type anything. The code points come from Unicode; the bytes from each encoding. The emoji needs two UTF-16 units, a surrogate pair, which is why JavaScript says "๐".length is 2.
Side by side
| ASCII | UTF-8 | UTF-16 | |
|---|---|---|---|
| Covers | 128 characters | All of Unicode | All of Unicode |
| Bytes per character | 1 | 1 to 4 | 2 or 4 |
| A | 41 | 41 | 00 41 |
| รฉ | not in ASCII | C3 A9 | 00 E9 |
| ๐ (U+1F600) | not in ASCII | F0 9F 98 80 | D8 3D DE 00 |
| ASCII text unchanged? | Yes | Yes | No: every character takes 2 bytes |
| Mostly used in | Old protocols, as a subset | The web, files, Linux, most APIs | JavaScript and Java strings, Windows APIs |
Why UTF-8 won
UTF-8 is backward compatible with ASCII: an ASCII file is already valid UTF-8, byte for byte. It has no byte-order question, it never produces a zero byte except for the character U+0000, and a program can always find where a character starts. Those properties made it the default for the web, for file formats such as JSON, and for most of the internet's protocols.
UTF-16 survives where it was built in early: JavaScript and Java strings, and Windows' own APIs, all designed when 65,536 characters looked like enough. When Unicode grew past that, characters above U+FFFF had to be split across two 16-bit units, a surrogate pair.
When it goes wrong
Text saved in one encoding and read in another turns into garbage such as รขโฌโข for an apostrophe. That's mojibake, and the mojibake fixer repairs the most common kind. The UTF-8 decoder explains bytes that aren't valid UTF-8 at all.
Questions
Is UTF-8 the same as Unicode?
No. Unicode is the list of characters and their numbers; UTF-8 is a way to store those numbers as bytes.
Is ASCII a subset of UTF-8?
Yes. The 128 ASCII characters have the same single-byte values in UTF-8.
Why is the length of an emoji 2 in JavaScript?
JavaScript strings are UTF-16, and characters above U+FFFF, emoji included, take two 16-bit units.
Should I use UTF-8 or UTF-16?
UTF-8 for files, the web and anything that leaves your program. UTF-16 mostly where a platform already uses it internally.
Sources
Added . What's new






