Database Replication
- In Turkish
- Veritabanı Replikasyonu
In short
Database replication is the continuous copying of data from one database server to others, so several servers hold the same data for reliability and scale.
What is database replication?
Replication keeps copies of the same database on more than one server. In the most common setup, one server, called the primary or leader, accepts all writes, and one or more replicas, also called followers or read replicas, receive a stream of those changes and apply them to their own copy. Older documentation calls this master-slave replication.
Replication serves three main purposes. It improves availability, because if the primary fails a replica can be promoted to take its place, a process called failover; it scales reads, because read-only queries can be spread across replicas; and it can place copies of the data closer to users in other regions. Most relational and NoSQL databases, including PostgreSQL, MySQL, MongoDB, and Cassandra, support it.
Replication can be synchronous or asynchronous. With synchronous replication, the primary waits for a replica to confirm each change before reporting success, which is safer but slower; with asynchronous replication, it doesn't wait, so replicas can lag slightly behind and a user might not see their own update if the next read goes to a lagging replica. Some systems use multi-leader or leaderless replication, where several nodes accept writes, at the cost of having to resolve conflicting updates.
It is a bit like a teacher's answer key photocopied for several assistants: any assistant can answer questions, but corrections are made on the original and then copied out again. Replication is often confused with backups and with sharding. A replica copies mistakes such as an accidental DELETE almost instantly, so it is not a backup, and unlike sharding, which splits different data across servers, replication gives each server the same data.
At a glance
Key takeaways
- Replication copies the same data to multiple database servers.
- In primary-replica setups, the primary handles writes and replicas serve reads.
- Failover promotes a replica if the primary goes down.
- Asynchronous replication is faster but lets replicas lag behind the primary.
- A replica is not a backup, because mistakes are replicated too.
Example
// Using node-postgres (pg) with two connection pools
import pg from "pg";
const primary = new pg.Pool({ connectionString: process.env.PRIMARY_DATABASE_URL });
const replica = new pg.Pool({ connectionString: process.env.REPLICA_DATABASE_URL });
await primary.query("UPDATE users SET name = $1 WHERE id = $2", ["Ada", 42]);
// With asynchronous replication, this read may briefly return the old name
const { rows } = await replica.query("SELECT name FROM users WHERE id = $1", [42]);Readers ask
What is the difference between replication and sharding?
Replication copies the same data to several servers, mainly for availability and read scaling. Sharding splits different parts of the data across servers to scale storage and writes, and each shard is often replicated as well.
What is replication lag?
Replication lag is the delay between a change being committed on the primary and that change appearing on a replica. It is usually milliseconds but can grow under heavy load, so reads that must see the latest data should go to the primary.
Is database replication the same as a backup?
No. Replication quickly copies every change, including accidental deletes and corrupted data, to the replicas. Backups are point-in-time snapshots that let you restore data from before a mistake, so you need both.
Often compared
See also
- ShardingDatabases, p. 39Sharding is a way of scaling a database by splitting its data across several servers, called shards, so each one stores and handles only part of the total.
- DatabaseDatabases, p. 6A database is an organized collection of data stored on a computer, managed by software that lets applications save, search, and update it efficiently.
- CAP TheoremDatabases, p. 2The CAP theorem says that if a network failure splits a distributed database, the system must choose between consistency and availability; it can't have both.
- TransactionDatabases, p. 47A transaction is a group of database operations that succeed or fail as a single unit, so the data is never left in a half-finished, inconsistent state.
- Load BalancerDevOps & Cloud, p. 34A load balancer is a server or service that spreads incoming traffic across several backend servers so no single one is overloaded and the app stays available.
- Consensus AlgorithmSoftware Architecture, p. 8A consensus algorithm lets a group of machines agree on a single value or an ordered log of decisions, even when some of them crash or messages are lost.
Spotted a mistake or something missing on this page?Suggest an edit