Skip to main content

Learning path · Beginner

Databases from scratch

Tables, queries and transactions, then what changes when data outgrows one server.

Start with tables and SQL, learn to design a schema that stays fast and correct, then see how databases spread across many machines.

31 pages4 chaptersabout 1 hours of reading

  • Databases

Not started yet0/31 read

Start with Database

Progress comes from your reading history, kept only in this browser.

Chapter 1Tables and queries

  1. 1DatabaseDatabases, p. 6A database is an organized collection of data stored on a computer, managed by software that lets applications save, search, and update it efficiently.
  2. 2Relational DatabaseDatabases, p. 38A relational database stores data in tables of rows and columns, links those tables through keys, and lets you query and combine the data with SQL.
  3. 3PostgreSQLDatabases, p. 35PostgreSQL is a free, open-source relational database known for reliability, strict standards support and extensions, and widely used for web applications.
  4. 4SQLDatabases, p. 40SQL is the standard language for working with relational databases, used to create tables and to insert, query, update, and delete the data stored in them.
  5. 5Database SchemaDatabases, p. 11A database schema is the blueprint of a database that defines its tables, columns, data types, relationships, and the rules that stored data must follow.
  6. 6Primary KeyDatabases, p. 36A primary key is a column, or set of columns, whose value uniquely identifies each row in a database table and can never be empty or duplicated.
  7. 7Foreign KeyDatabases, p. 21A foreign key is a column in one database table that refers to the primary key of another table, linking related rows and keeping those references valid.
  8. 8SQL JOINDatabases, p. 41A SQL JOIN is a query operation that combines rows from two or more tables into one result, matching them on related columns such as a foreign key.

Chapter 2Designing well

  1. 9Database NormalizationDatabases, p. 9Database normalization is the process of organizing tables so each fact is stored only once, reducing duplicate data and preventing inconsistent updates.
  2. 10DenormalizationDatabases, p. 14Denormalization is the deliberate duplication of data across tables or documents so that frequent reads need fewer joins, at the cost of more complex writes.
  3. 11Database IndexDatabases, p. 7A database index is a data structure that helps a database find rows quickly without scanning a whole table, much like the index at the back of a book.
  4. 12N+1 Query ProblemDatabases, p. 28The N+1 query problem is a performance bug where code runs one query to load a list and then one extra query per item, instead of fetching it all at once.
  5. 13Database ViewDatabases, p. 13A database view is a saved SQL query that behaves like a virtual table, so you can select from it by name instead of repeating the underlying query each time.

Chapter 3Keeping data correct

  1. 14TransactionDatabases, p. 47A transaction is a group of database operations that succeed or fail as a single unit, so the data is never left in a half-finished, inconsistent state.
  2. 15OLTPDatabases, p. 31OLTP (online transaction processing) describes databases built for many small, fast reads and writes from everyday operations, such as placing orders.
  3. 16ACIDDatabases, p. 1ACID is a set of four guarantees, atomicity, consistency, isolation, and durability, that keep database transactions reliable even when errors or crashes occur.
  4. 17Isolation LevelDatabases, p. 24An isolation level is a database setting that controls how much concurrent transactions can see of each other's changes, trading strictness for speed.
  5. 18Optimistic LockingDatabases, p. 32Optimistic locking is a concurrency technique that lets transactions proceed without holding locks and checks a version number at save time to detect conflicts.
  6. 19Eventual ConsistencyDatabases, p. 19Eventual consistency is a guarantee that, if no new updates are made, all copies of a piece of data in a distributed system will become identical over time.
  7. 20CAP TheoremDatabases, p. 2The CAP theorem says that if a network failure splits a distributed database, the system must choose between consistency and availability; it can't have both.

Chapter 4Beyond one server

  1. 21NoSQLDatabases, p. 29NoSQL is a family of databases that store data in models other than relational tables, such as documents, key-value pairs, wide columns, or graphs.
  2. 22Document DatabaseDatabases, p. 15A document database is a NoSQL database that stores each record as a self-contained document, usually JSON-like, whose fields can differ from record to record.
  3. 23MongoDBDatabases, p. 26MongoDB is a document database that stores data as flexible JSON-like documents instead of table rows, so records in one collection can have different fields.
  4. 24Key-Value StoreDatabases, p. 25A key-value store is a NoSQL database that saves each piece of data under a unique key, so an application can read or write it by that key very quickly.
  5. 25RedisDatabases, p. 37Redis is an in-memory key-value store that reads and writes in well under a millisecond, which makes it a popular cache, session store and message broker.
  6. 26Database ReplicationDatabases, p. 10Database replication is the continuous copying of data from one database server to others, so several servers hold the same data for reliability and scale.
  7. 27PartitioningDatabases, p. 34Partitioning splits a large table into smaller partitions by a rule such as date ranges, so queries can skip irrelevant data and old data is easy to remove.
  8. 28ShardingDatabases, p. 39Sharding is a way of scaling a database by splitting its data across several servers, called shards, so each one stores and handles only part of the total.
  9. 29Data WarehouseDatabases, p. 5A data warehouse is a central database built for analytics that collects historical data from many sources so teams can run large reporting queries quickly.
  10. 30OLAPDatabases, p. 30OLAP (online analytical processing) describes systems built to answer complex analytical questions over large amounts of historical data quickly.
  11. 31Data LakeDatabases, p. 4A data lake is a central storage repository that holds large amounts of raw data in its original format, structured or not, until someone needs to analyze it.

Along the way, compare

Pairs on this path that are easy to mix up, side by side.

More

Settings