Distributed System
- In Turkish
- Dağıtık Sistem
- Pronunciation
- dih-STRIB-yoo-tid SIS-tum
In short
A distributed system is a set of computers that work together over a network and appear to their users as a single system.
What is a distributed system?
Systems are distributed for scale, reliability and reach. One machine can only handle so much traffic and data, and it will eventually fail. Spreading the work across many machines lets a service grow, keep running when some of them break, and serve users from data centers close to them. Search engines, social networks, cloud databases, payment systems and microservice architectures are all distributed systems.
Distribution brings problems a single computer never has. Messages can be delayed, lost or duplicated; one machine can crash while others keep running, a situation called partial failure; clocks on different machines disagree; and the network can split so that groups of machines can't reach each other. Designs must decide how to keep data consistent and how to agree on decisions, using replication, consensus algorithms, idempotent operations and retries.
In the 1990s, engineers at Sun Microsystems listed the fallacies of distributed computing, assumptions developers make that turn out to be false: the network is reliable, latency is zero, bandwidth is infinite, the network is secure, the topology doesn't change, there is one administrator, transport costs nothing and the network is the same everywhere. Every one of them causes real outages.
A common misconception is that distributing a system automatically makes it more reliable. More machines means more things that can fail, and more ways for them to fail together. Reliability comes from deliberate design: timeouts, retries with backoff, circuit breakers, redundancy without single points of failure, and observability to understand what is happening.
Key takeaways
- A distributed system is many networked computers acting as one.
- It is built for scale, fault tolerance and serving users near them.
- Partial failures, lost messages, clock drift and network splits are normal.
- The fallacies of distributed computing list assumptions that break systems.
- Reliability requires deliberate design, not just more machines.
Readers ask
What are examples of distributed systems?
Web applications running on many servers behind a load balancer, cloud databases such as DynamoDB, streaming platforms such as Kafka, content delivery networks, blockchains and microservice architectures.
Why are distributed systems hard?
Because components fail independently and communicate over an unreliable network. Code must handle delays, lost or duplicated messages, inconsistent data and machines that disagree, all while staying correct.
What is the CAP theorem's role in distributed systems?
It states that when the network partitions, a distributed data store must choose between consistency and availability. It is one of the basic trade-offs designers of distributed systems face.
See also
- CAP TheoremDatabases, p. 2The CAP theorem says that if a network failure splits a distributed database, the system must choose between consistency and availability; it can't have both.
- MicroservicesSoftware Architecture, p. 27Microservices are an architectural style where an application is split into small, independently deployable services that communicate over a network.
- Database ReplicationDatabases, p. 10Database replication is the continuous copying of data from one database server to others, so several servers hold the same data for reliability and scale.
- Consensus AlgorithmSoftware Architecture, p. 8A consensus algorithm lets a group of machines agree on a single value or an ordered log of decisions, even when some of them crash or messages are lost.
- Fault ToleranceSoftware Architecture, p. 20Fault tolerance is the ability of a system to keep working correctly, perhaps at reduced capacity, when some of its hardware or software components fail.
- Eventual ConsistencyDatabases, p. 19Eventual consistency is a guarantee that, if no new updates are made, all copies of a piece of data in a distributed system will become identical over time.
Spotted a mistake or something missing on this page?Suggest an edit