Character Encoding
- In Turkish
- Karakter kodlaması
In short
Character encoding is the set of rules that turns text into bytes and back again, so any letter can be stored and sent; today the standard is UTF-8.
What is character encoding?
Computers store only numbers, so text has to become bytes before it can be saved to a file or sent over a network. A character encoding defines how. Unicode gives every character in every writing system its own number, called a code point: the letter A is U+0041 and the Turkish ç is U+00E7. An encoding such as UTF-8 then says how each code point is written as bytes.
ASCII, from the 1960s, used 7 bits and covered 128 characters: English letters, digits, punctuation and control codes. Dozens of incompatible 8-bit encodings followed, one per region, such as ISO-8859-9 for Turkish. UTF-8 ended that by encoding all of Unicode in 1 to 4 bytes per character while leaving plain ASCII text unchanged; today almost every web page uses it, and it is the default in most languages and tools.
Text reads correctly only if the reader uses the same encoding as the writer. Open a UTF-8 file as Windows-1254 and ç turns into ç, a garbling known as mojibake. That is why web pages declare <meta charset="utf-8">, HTTP responses name a charset in the Content-Type header, and databases and source files are best set to UTF-8 from the start.
Key takeaways
- A character encoding maps text to bytes and back.
- Unicode numbers every character; an encoding such as UTF-8 decides how those numbers become bytes.
- UTF-8 uses 1 to 4 bytes per character and leaves ASCII text unchanged.
- Reading text with the wrong encoding garbles it, so use UTF-8 everywhere.
Example
const bytes = new TextEncoder().encode("Aç"); // TextEncoder always writes UTF-8
console.log(bytes); // Uint8Array(3) [ 65, 195, 167 ]
console.log("ç".codePointAt(0).toString(16)); // "e7": the code point U+00E7
// The same bytes read as Windows-1254 instead of UTF-8:
console.log(new TextDecoder("windows-1254").decode(bytes)); // "Aç"Readers ask
What is the difference between Unicode and UTF-8?
Unicode is the catalog: it gives each character a number. UTF-8 is one way to write those numbers as bytes; UTF-16 and UTF-32 are others. So a text can be in Unicode while its file is encoded as UTF-8.
Why do I see characters like ç instead of ç?
The bytes were written in one encoding and read in another, most often UTF-8 read as a one-byte encoding such as Windows-1252 or Windows-1254. Read the text with the encoding it was written in, and use UTF-8 throughout so the question doesn't come up.
See also
- StringProgramming Fundamentals, p. 54A string is a data type that represents text as an ordered sequence of characters, such as a name, a sentence, a URL, or the contents of a file.
- Data TypeProgramming Fundamentals, p. 14A data type is a classification that tells a program what kind of value a piece of data holds, such as a number or text, and which operations work on it.
- Base64Programming Fundamentals, p. 5Base64 is a way to write any binary data, such as an image or a key, using only 64 safe text characters, so it can pass through systems built for text.
- URLWeb Development, p. 52A URL is the address of a resource on the web, made of parts such as a scheme, a host, a path and a query string that tell a browser where and how to fetch it.
- HTTPWeb Development, p. 19HTTP is the protocol that browsers, apps, and servers use to exchange web pages and data through a simple cycle of requests and responses.
- HTMLWeb Development, p. 18HTML is the markup language that defines the structure and content of web pages, such as headings, paragraphs, links, images, and forms.
Sources
Spotted a mistake or something missing on this page?Suggest an edit