Skip to main content

Character Encoding

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/character-encoding

In short

Character encoding is the set of rules that turns text into bytes and back again, so any letter can be stored and sent; today the standard is UTF-8.

What is character encoding?

Computers store only numbers, so text has to become bytes before it can be saved to a file or sent over a network. A character encoding defines how. Unicode gives every character in every writing system its own number, called a code point: the letter A is U+0041 and the Turkish ç is U+00E7. An encoding such as UTF-8 then says how each code point is written as bytes.

ASCII, from the 1960s, used 7 bits and covered 128 characters: English letters, digits, punctuation and control codes. Dozens of incompatible 8-bit encodings followed, one per region, such as ISO-8859-9 for Turkish. UTF-8 ended that by encoding all of Unicode in 1 to 4 bytes per character while leaving plain ASCII text unchanged; today almost every web page uses it, and it is the default in most languages and tools.

Text reads correctly only if the reader uses the same encoding as the writer. Open a UTF-8 file as Windows-1254 and ç turns into ç, a garbling known as mojibake. That is why web pages declare <meta charset="utf-8">, HTTP responses name a charset in the Content-Type header, and databases and source files are best set to UTF-8 from the start.

Key takeaways

  • A character encoding maps text to bytes and back.
  • Unicode numbers every character; an encoding such as UTF-8 decides how those numbers become bytes.
  • UTF-8 uses 1 to 4 bytes per character and leaves ASCII text unchanged.
  • Reading text with the wrong encoding garbles it, so use UTF-8 everywhere.

Example

From characters to UTF-8 bytes, and back with the wrong encodingjavascript
const bytes = new TextEncoder().encode("Aç"); // TextEncoder always writes UTF-8
console.log(bytes);                              // Uint8Array(3) [ 65, 195, 167 ]
console.log("ç".codePointAt(0).toString(16));    // "e7": the code point U+00E7

// The same bytes read as Windows-1254 instead of UTF-8:
console.log(new TextDecoder("windows-1254").decode(bytes)); // "Aç"

Readers ask

What is the difference between Unicode and UTF-8?

Unicode is the catalog: it gives each character a number. UTF-8 is one way to write those numbers as bytes; UTF-16 and UTF-32 are others. So a text can be in Unicode while its file is encoded as UTF-8.

Why do I see characters like ç instead of ç?

The bytes were written in one encoding and read in another, most often UTF-8 read as a one-byte encoding such as Windows-1252 or Windows-1254. Read the text with the encoding it was written in, and use UTF-8 throughout so the question doesn't come up.

See also

Sources

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings