HTML Character Sets
HTML Character Sets
To correctly display an HTML page, the browser must know which character set to use.
The character set used in the early days of the World Wide Web was ASCII. ASCII supports the numbers 0-9, uppercase and lowercase English alphabets, and some special characters.
Complete ASCII Reference Manual。
Since many countries use characters that do not belong to ASCII, the default character set of modern browsers is ISO-8859-1.
Complete ISO-8859-1 Reference Manual。
If a web page uses a character set other than ISO-8859-1, it should be specified in the <meta> tag.
ISO Character Sets
ISO character sets are standard character sets defined by the International Organization for Standardization (ISO) for different alphabets/languages.
The following lists the different character sets used around the world:
| Character Set | Description | Usage Range |
|---|---|---|
| ISO-8859-1 | Latin alphabet part 1 | North America, Western Europe, Latin America, Caribbean, Canada, Africa |
| ISO-8859-2 | Latin alphabet part 2 | Eastern Europe |
| ISO-8859-3 | Latin alphabet part 3 | SE Europe, Esperanto, other miscellaneous |
| ISO-8859-4 | Latin alphabet part 4 | Scandinavia/Baltic (and other parts not included in ISO-8859-1) |
| ISO-8859-5 | Latin/Cyrillic part 5 | Languages using the Old Slavic alphabet, such as Bulgarian, Belarusian, Russian, Macedonian |
| ISO-8859-6 | Latin/Arabic part 6 | Languages using the Arabic alphabet |
| ISO-8859-7 | Latin/Greek part 7 | Modern Greek, and mathematical symbols derived from Greek |
| ISO-8859-8 | Latin/Hebrew part 8 | Languages using Hebrew |
| ISO-8859-9 | Latin 5 part 9 | Turkish. Same as ISO-8859-1, except that Turkish characters replace Icelandic characters. |
| ISO-8859-10 | Latin 6 | Lappish, Germanic, Eskimo-Nordic languages |
| ISO-8859-15 | Latin 9 (aka Latin 0) | Similar to ISO 8859-1, the euro sign and some other characters replace some less commonly used symbols. |
| ISO-2022-JP | Latin/Japanese part 1 | Japanese |
| ISO-2022-JP-2 | Latin/Japanese part 2 | Japanese |
| ISO-2022-KR | Latin/Korean part 1 | Korean |
Unicode Standard
Because the character sets listed above have capacity limitations and are incompatible with multilingual environments, the Unicode Consortium developed the Unicode Standard.
The Unicode Standard covers all characters, punctuation, and symbols in the world.
Regardless of platform, program, or language, Unicode can process, store, and exchange text data.
Unicode Consortium
The Unicode Consortium developed the Unicode Standard. Their goal is to replace existing character sets with the standard Unicode Transformation Format (UTF).
The Unicode Standard has been successful. Unicode is implemented in XML, Java, ECMAScript (JavaScript), LDAP, CORBA 3.0, and WML. Unicode is also supported in many operating systems and all modern browsers.
The Unicode Consortium cooperates with leading standard development organizations, such as ISO, W3C, and ECMA.
Unicode can be compatible with different character sets. The most common encoding methods are UTF-8 and UTF-16:
| Character Set | Description |
|---|---|
| UTF-8 | Characters in UTF-8 can be 1-4 bytes long. UTF-8 can represent any character in the Unicode Standard. UTF-8 is backward compatible with ASCII. UTF-8 is the preferred encoding for web pages and email. |
| UTF-16 | The 16-bit Unicode Transformation Format is a Unicode variable character encoding that can encode all Unicode code points. UTF-16 is mainly used in operating systems and environments, such as Microsoft Windows 2000/XP/2003/Vista/CE, and Java and .NET bytecode environments. |
Tip:The first 256 Unicode character set characters correspond to the 256 ISO-8859-1 characters.
Tip:All HTML 4 browsers already support UTF-8, and all XHTML and XML processors support UTF-8 and UTF-16!
Other Extensions