HTML Character Sets


HTML Character Sets

To correctly display an HTML page, the browser must know which character set to use.

The character set used in the early days of the World Wide Web was ASCII. ASCII supports the numbers 0-9, uppercase and lowercase English alphabets, and some special characters.

Complete ASCII Reference Manual。

Since many countries use characters that do not belong to ASCII, the default character set of modern browsers is ISO-8859-1.

Complete ISO-8859-1 Reference Manual。

If a web page uses a character set other than ISO-8859-1, it should be specified in the <meta> tag.

Try it yourself


ISO Character Sets

ISO character sets are standard character sets defined by the International Organization for Standardization (ISO) for different alphabets/languages.

The following lists the different character sets used around the world:

Character Set Description Usage Range
ISO-8859-1 Latin alphabet part 1 North America, Western Europe, Latin America, Caribbean, Canada, Africa
ISO-8859-2 Latin alphabet part 2 Eastern Europe
ISO-8859-3 Latin alphabet part 3 SE Europe, Esperanto, other miscellaneous
ISO-8859-4 Latin alphabet part 4 Scandinavia/Baltic (and other parts not included in ISO-8859-1)
ISO-8859-5 Latin/Cyrillic part 5 Languages using the Old Slavic alphabet, such as Bulgarian, Belarusian, Russian, Macedonian
ISO-8859-6 Latin/Arabic part 6 Languages using the Arabic alphabet
ISO-8859-7 Latin/Greek part 7 Modern Greek, and mathematical symbols derived from Greek
ISO-8859-8 Latin/Hebrew part 8 Languages using Hebrew
ISO-8859-9 Latin 5 part 9 Turkish. Same as ISO-8859-1, except that Turkish characters replace Icelandic characters.
ISO-8859-10 Latin 6 Lappish, Germanic, Eskimo-Nordic languages
ISO-8859-15 Latin 9 (aka Latin 0) Similar to ISO 8859-1, the euro sign and some other characters replace some less commonly used symbols.
ISO-2022-JP Latin/Japanese part 1 Japanese
ISO-2022-JP-2 Latin/Japanese part 2 Japanese
ISO-2022-KR Latin/Korean part 1 Korean


Unicode Standard

Because the character sets listed above have capacity limitations and are incompatible with multilingual environments, the Unicode Consortium developed the Unicode Standard.

The Unicode Standard covers all characters, punctuation, and symbols in the world.

Regardless of platform, program, or language, Unicode can process, store, and exchange text data.


Unicode Consortium

The Unicode Consortium developed the Unicode Standard. Their goal is to replace existing character sets with the standard Unicode Transformation Format (UTF).

The Unicode Standard has been successful. Unicode is implemented in XML, Java, ECMAScript (JavaScript), LDAP, CORBA 3.0, and WML. Unicode is also supported in many operating systems and all modern browsers.

The Unicode Consortium cooperates with leading standard development organizations, such as ISO, W3C, and ECMA.

Unicode can be compatible with different character sets. The most common encoding methods are UTF-8 and UTF-16:

Character Set Description
UTF-8 Characters in UTF-8 can be 1-4 bytes long. UTF-8 can represent any character in the Unicode Standard. UTF-8 is backward compatible with ASCII. UTF-8 is the preferred encoding for web pages and email.
UTF-16 The 16-bit Unicode Transformation Format is a Unicode variable character encoding that can encode all Unicode code points. UTF-16 is mainly used in operating systems and environments, such as Microsoft Windows 2000/XP/2003/Vista/CE, and Java and .NET bytecode environments.

Tip:The first 256 Unicode character set characters correspond to the 256 ISO-8859-1 characters.

Tip:All HTML 4 browsers already support UTF-8, and all XHTML and XML processors support UTF-8 and UTF-16!

Other Extensions