HTML Character Sets
A character set defines the characters that can be represented in a document, while a character encoding defines how those characters are stored as bytes. Modern HTML documents should use UTF-8, which supports characters from languages and writing systems around the world.
Declaring the correct character encoding helps browsers display text, symbols, punctuation, emoji, and other Unicode characters correctly.
Character Encoding
Computers store text as numbers. A character encoding provides the rules used to convert those numbers into the characters displayed on a webpage.
If a document is interpreted using the wrong encoding, characters may be displayed incorrectly. For example, quotation marks, accented letters, currency symbols, or other characters may appear as unexpected symbols.
UTF-8
UTF-8 is the standard character encoding for HTML documents. It can represent every Unicode character while remaining compatible with the ASCII characters commonly used in HTML source code.
UTF-8 supports Latin letters, accented characters, Greek, Cyrillic, Arabic, Asian writing systems, mathematical symbols, emoji, and thousands of other characters.
| Feature | Description |
|---|---|
| Unicode Support | Can represent the complete range of Unicode characters. |
| ASCII Compatibility | The first 128 characters correspond to ASCII. |
| Variable Length | Characters are encoded using one to four bytes. |
| Web Standard | UTF-8 is the encoding used for modern HTML documents. |
Declaring the Character Encoding
Use the <meta charset="utf-8"> element near the beginning of the document's <head> to declare UTF-8 encoding.
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<title>HTML Example</title>
</head>
<body>
<h1>Hello, World!</h1>
</body>
</html>
The character encoding declaration should appear as early as possible in the document and must be completely contained within the first 1024 bytes of the HTML file.
Unicode
Unicode is a character standard that assigns a unique code point to characters used by writing systems, symbols, punctuation, emoji, and other forms of text.
Unicode code points are commonly written using U+ followed by a hexadecimal value. For example, the copyright sign is Unicode code point U+00A9, while the Greek capital letter Omega is U+03A9.
| Character | Unicode | Description |
|---|---|---|
| A | U+0041 | Latin capital letter A |
| © | U+00A9 | Copyright sign |
| Ω | U+03A9 | Greek capital letter Omega |
| € | U+20AC | Euro sign |
| ✓ | U+2713 | Check mark |
| 😀 | U+1F600 | Grinning face |
ASCII
ASCII is an older character encoding standard containing 128 characters. It includes uppercase and lowercase English letters, digits, punctuation marks, and control characters.
UTF-8 preserves the same values for these first 128 characters, which makes ASCII text compatible with UTF-8.
| Character | Decimal | Hexadecimal | Description |
|---|---|---|---|
| A | 65 | 41 | Uppercase A |
| a | 97 | 61 | Lowercase a |
| 0 | 48 | 30 | Digit zero |
| & | 38 | 26 | Ampersand |
| < | 60 | 3C | Less-than sign |
| > | 62 | 3E | Greater-than sign |
HTML Character References
HTML character references provide another way to represent characters in HTML source code. A character can be written using a named character reference, a decimal numeric reference, or a hexadecimal numeric reference.
| Type | Example | Result |
|---|---|---|
| Named | &copy; | © |
| Decimal | © | © |
| Hexadecimal | © | © |
Character references are especially useful for characters that have special meaning in HTML syntax, characters that are difficult to type, or when the source code should explicitly identify a particular character value.
Common Character Encodings
Older webpages and documents may use character encodings other than UTF-8. These encodings may still be encountered when maintaining or converting legacy content.
| Encoding | Description |
|---|---|
UTF-8 | Unicode encoding used for modern HTML and recommended for new webpages. |
Windows-1252 | Legacy single-byte encoding commonly used by older Western-language webpages. |
ISO-8859-1 | Older single-byte character encoding for Western European languages. |
Shift_JIS | Legacy encoding used for Japanese text. |
GBK | Legacy encoding used for Simplified Chinese text. |
Big5 | Legacy encoding used for Traditional Chinese text. |
For new HTML documents, use UTF-8 rather than choosing a legacy character encoding.
Character Encoding Best Practices
- Use UTF-8 for new HTML documents.
- Include
<meta charset="utf-8">near the beginning of the<head>. - Save the HTML file itself using UTF-8 encoding.
- Make sure the HTTP character encoding, when supplied by the server, agrees with the document encoding.
- Use HTML character references when needed for HTML syntax or when they make the source easier to understand.
- Do not rely on the browser to guess an undeclared character encoding.
