Serialization
- In Turkish
- Serileştirme
In short
Serialization is the process of converting in-memory data structures into a format such as JSON or bytes, so they can be stored or sent over a network.
What is serialization?
Serialization turns data that lives in a program's memory, such as an object, a list, or a nested structure, into a sequence of characters or bytes that can be saved to a file, stored in a database or cache, or sent over a network. The reverse process, rebuilding the in-memory data from that format, is called deserialization. Together they let different programs, machines, and even programming languages share the same data.
It is needed because in-memory data contains details that only make sense inside one running program, such as memory addresses and references between objects. A serializer walks through the data and writes out its values in an agreed format. Text formats like JSON, XML, YAML, and CSV are human-readable, while binary formats like Protocol Buffers, MessagePack, and Avro are smaller and faster to process, and some languages have native formats such as Python's pickle.
A good analogy is flat-pack furniture: a desk is taken apart and packed into a flat box for shipping, then assembled again at its destination by following the instructions. Serialization happens every time an API returns JSON, a web app saves state in local storage, a message is placed on a queue, or a cache stores an object, and it is often a hidden cost behind slow requests.
Serialization is often confused with encoding and encryption. Serialization decides how the structure of data is written out, encoding such as UTF-8 or Base64 decides how characters or bytes are represented, and encryption hides data from anyone without the key. Deserializing untrusted input with native formats like pickle or Java's built-in serialization is dangerous, because a crafted payload can run code, so outside data should use a data-only format like JSON and be validated.
Key takeaways
- Serialization converts in-memory data into text or bytes for storage or transfer.
- Deserialization rebuilds the in-memory data from that format.
- JSON, XML, and YAML are text formats; Protocol Buffers and MessagePack are binary.
- APIs, caches, message queues, and files all depend on serialization.
- Never deserialize untrusted data with formats that can run code, such as
pickle.
Example
import json
from datetime import date
user = {"id": 42, "name": "Ada", "roles": ["admin"], "joined": date(2026, 9, 30)}
# Serialize: Python dict -> JSON text (the date must be converted to a string)
text = json.dumps(user, default=str)
print(text) # {"id": 42, "name": "Ada", "roles": ["admin"], "joined": "2026-09-30"}
# Deserialize: JSON text -> Python dict
restored = json.loads(text)
print(restored["name"]) # Ada
print(type(restored["joined"])) # <class 'str'>: the original date type is lostReaders ask
What is the difference between serialization and deserialization?
Serialization converts in-memory data into a format that can be stored or sent, such as a JSON string. Deserialization does the opposite, reading that format and rebuilding the data structures in memory.
Is JSON serialization?
JSON is a data format, and converting data to JSON is one of the most common kinds of serialization. In JavaScript, JSON.stringify() serializes a value and JSON.parse() deserializes it.
Why is insecure deserialization dangerous?
Some native formats can recreate any type of object and trigger code while doing so, so an attacker who controls the input may be able to run commands on the server. Use data-only formats such as JSON for untrusted input and validate the result.
See also
- JSONBackend & APIs, p. 25JSON is a lightweight, text-based format for storing and exchanging structured data as key-value pairs and lists, readable by both humans and machines.
- APIBackend & APIs, p. 2An API is a set of rules that lets one piece of software request data or actions from another in a predictable, documented way.
- Message QueueBackend & APIs, p. 29A message queue is a component that stores messages from one service until another is ready to process them, so parts of a system can work asynchronously.
- CacheBackend & APIs, p. 8A cache is a fast, temporary storage layer that keeps copies of frequently used data so later requests can be served quickly without repeating slow work.
- gRPCBackend & APIs, p. 20gRPC is an open-source framework for calling functions on a remote server as if they were local, using Protocol Buffers and HTTP/2 for fast, typed messages.
- YAMLDevOps & Cloud, p. 54YAML is a human-readable data format that uses indentation instead of brackets, widely used for configuration files in DevOps tools and CI/CD pipelines.
Spotted a mistake or something missing on this page?Suggest an edit