StringMash.com

ASCII vs Unicode vs UTF-8

A character set, a catalogue, and two ways to write the catalogue down.

The short answer

ASCII is a set of 128 characters, the English letters, digits and punctuation, each stored in one byte. Unicode is a catalogue that gives a number, a code point, to every character in every script, more than 150,000 of them. UTF-8 and UTF-16 are encodings: ways of storing those code points as bytes. UTF-8 uses one to four bytes per character and is identical to ASCII for ASCII text; UTF-16 uses two or four.

Look inside any character

Type anything. The code points come from Unicode; the bytes from each encoding. The emoji needs two UTF-16 units, a surrogate pair, which is why JavaScript says "๐Ÿ˜€".length is 2.

Unicode code points
U+0041 U+00E9 U+1F60021 characters, 21 bytes
UTF-8 bytes
41 C3 A9 F0 9F 98 8020 characters, 20 bytes
Open the UTF-8 encoder, with a per-character breakdown โ†’

Side by side

ASCIIUTF-8UTF-16
Covers128 charactersAll of UnicodeAll of Unicode
Bytes per character11 to 42 or 4
A414100 41
รฉnot in ASCIIC3 A900 E9
๐Ÿ˜€ (U+1F600)not in ASCIIF0 9F 98 80D8 3D DE 00
ASCII text unchanged?YesYesNo: every character takes 2 bytes
Mostly used inOld protocols, as a subsetThe web, files, Linux, most APIsJavaScript and Java strings, Windows APIs

Why UTF-8 won

UTF-8 is backward compatible with ASCII: an ASCII file is already valid UTF-8, byte for byte. It has no byte-order question, it never produces a zero byte except for the character U+0000, and a program can always find where a character starts. Those properties made it the default for the web, for file formats such as JSON, and for most of the internet's protocols.

UTF-16 survives where it was built in early: JavaScript and Java strings, and Windows' own APIs, all designed when 65,536 characters looked like enough. When Unicode grew past that, characters above U+FFFF had to be split across two 16-bit units, a surrogate pair.

When it goes wrong

Text saved in one encoding and read in another turns into garbage such as รขโ‚ฌโ„ข for an apostrophe. That's mojibake, and the mojibake fixer repairs the most common kind. The UTF-8 decoder explains bytes that aren't valid UTF-8 at all.

Questions

Is UTF-8 the same as Unicode?

No. Unicode is the list of characters and their numbers; UTF-8 is a way to store those numbers as bytes.

Is ASCII a subset of UTF-8?

Yes. The 128 ASCII characters have the same single-byte values in UTF-8.

Why is the length of an emoji 2 in JavaScript?

JavaScript strings are UTF-16, and characters above U+FFFF, emoji included, take two 16-bit units.

Should I use UTF-8 or UTF-16?

UTF-8 for files, the web and anything that leaves your program. UTF-16 mostly where a platform already uses it internally.

Sources

Added . What's new