HTML Character Set


To correctly display an HTML page, the browser must know which character set (character encoding) to use.


HTML Character Set

What is the correct character encoding in HTML?

The default character encoding in HTML5 is UTF-8.

This has not always been the case. The character encoding for the early web was ASCII.

Later, from HTML 2.0 to HTML 4.01, ISO-8859-1 was recognized as the standard.

With the emergence of XML and HTML5, UTF-8 finally arrived, solving many character encoding problems.

Below is a brief overview of character encoding standards.


In the beginning: ASCII

Computer information (numbers, text, pictures) is stored electronically using binary 1s and 0s (01000101).

To standardize the storage of alphanumeric characters, ASCII (full name: American Standard Code for Information Interchange) was created. It defines a unique binary 7-bit number for each stored character, supporting numbers 0-9, uppercase/lowercase English letters (a-z, A-Z), and some special characters such as ! $ + - ( ) @ < >.

Because ASCII uses one byte (7 bits for the character, 1 bit for transmission parity control), it can only represent 128 different characters. Of these characters, 32 are reserved for other control purposes.

The biggest disadvantage of ASCII is that it excludes non-English letters.

ASCII is still widely used today, especially in large computer systems.

For more information on ASCII, please refer toComplete ASCII Reference Manual。


In Windows: ANSI

ANSI (also known as Windows-1252) is the default character set in Windows 95 and earlier Windows systems.

ANSI is an extension of ASCII, adding international characters. It uses a full byte (8 bits) to represent 256 different characters.

Since ANSI became the default character set in Windows, all browsers support ANSI.

For more information on ANSI, please refer toComplete ANSI Reference Manual。


In HTML 4: ISO-8859-1

Because most countries use characters outside ASCII, the default character encoding was changed to ISO-8859-1 in the HTML 2.0 standard.

ISO-8859-1 is an extension of ASCII, adding international characters. Like ANSI, it uses a full byte (8 bits) to represent 256 different characters.

Note When a browser detects ISO-8859-1 in a web page, it usually defaults to ANSI, because ANSI is basically equivalent to ISO-8859-1 except that ANSI has 32 additional characters.

If an HTML 4 web page uses a character set different from ISO-8859-1, it needs to be specified in the <meta> tag, as shown below:

Examples

<meta http-equiv="Content-Type" content="text/html;charset=ISO-8859-8">

Note

The default character set in HTML5 is UTF-8.
All HTML 4 processors support UTF-8, and all HTML5 and XML processors support UTF-8 and UTF-16.

For more information on ISO-8859-1, please refer toComplete ISO-8859-1 Reference Manual。


In HTML5: Unicode (UTF-8)

Since the character sets listed above are limited and incompatible in multilingual environments, the Unicode Consortium developed the Unicode Standard.

The Unicode Standard covers (almost) all characters, punctuation, and symbols.

Unicode enables the processing, storage, and transport of text to be independent of platform and language.

The default character encoding in HTML5 is UTF-8.

For more information on Unicode (UTF-8), seeComplete Unicode Reference Manual。


Other extensions