Data Lake
In short
A data lake is a central storage repository that holds large amounts of raw data in its original format, structured or not, until someone needs to analyze it.
What is a data lake?
A data lake is a large, central store for raw data in whatever form it arrives: database exports, application logs, clickstream events, JSON files, images, and sensor readings. The data is kept as files, usually in inexpensive and highly scalable object storage. Unlike a data warehouse, a data lake does not require you to define a schema before you store something.
Files in a lake are organized into folders, often partitioned by date or source, and are commonly saved in columnar formats such as Parquet that make analytical queries efficient. A catalog records which datasets exist and what their columns look like, and query engines read the files directly with SQL, applying the structure at read time, an approach called schema-on-read. Open table formats such as Apache Iceberg and Delta Lake add transactions, schema changes, and time travel on top of plain files, which is the basis of the lakehouse approach that combines lake storage with warehouse-style tables.
If a data warehouse is a supermarket with labeled shelves of ready-to-use products, a data lake is a reservoir where water from many streams collects in its natural state and is filtered only when someone draws from it. Data lakes are used for machine learning training data, data science exploration, cheap long-term retention, and log archives, especially when nobody knows yet exactly which questions the data will answer.
The key confusion is with a data warehouse. A warehouse stores cleaned, structured, modeled data designed for business reporting and applies the schema when data is written, while a lake stores everything raw and is cheaper and more flexible but slower to query and harder to trust. Without a catalog, quality checks, and access control, a lake turns into a data swamp that nobody can navigate. Many organizations use both, with ETL or ELT pipelines refining lake data into warehouse tables.
Key takeaways
- A data lake stores raw data of any type, usually as files in object storage.
- Structure is applied when the data is read, not when it is written.
- Columnar formats such as Parquet and a data catalog make lakes usable.
- A lakehouse adds warehouse-style tables and transactions on top of a lake.
- Without governance, a data lake becomes an unusable data swamp.
Example
-- Files in the lake are partitioned by date:
-- lake/clickstream/date=2026-09-29/part-0001.parquet
-- lake/clickstream/date=2026-09-30/part-0001.parquet
-- The engine reads the schema from the files themselves
SELECT page, COUNT(*) AS views
FROM read_parquet('lake/clickstream/date=2026-09-30/*.parquet')
GROUP BY page
ORDER BY views DESC
LIMIT 10;Readers ask
What is the difference between a data lake and a data warehouse?
A data lake stores raw data of any type cheaply and applies structure only when it is read. A data warehouse stores cleaned, structured data with a schema defined up front, which makes it faster and more reliable for business reporting.
What is a data lakehouse?
A lakehouse keeps data in open file formats on lake storage but adds a table layer with transactions, schemas, and fast SQL queries. It aims to give warehouse-style reliability without copying data into a separate warehouse.
What is a data swamp?
A data swamp is a data lake that has become disorganized, with undocumented, duplicated, or low-quality data that nobody can find or trust. Catalogs, ownership, and quality checks prevent it.
Often compared
See also
- Data WarehouseDatabases, p. 5A data warehouse is a central database built for analytics that collects historical data from many sources so teams can run large reporting queries quickly.
- ETLDatabases, p. 18ETL is a data integration process that extracts data from source systems, transforms it into a clean, consistent shape, and loads it into a target store.
- Object StorageDevOps & Cloud, p. 37Object storage is a way of storing data as whole objects, each with a unique key and metadata, in flat buckets that scale to huge numbers of files over HTTP.
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Database SchemaDatabases, p. 11A database schema is the blueprint of a database that defines its tables, columns, data types, relationships, and the rules that stored data must follow.
- Training DataAI & Machine Learning, p. 48Training data is the set of examples a machine learning model learns from, and its quality, size, and coverage largely determine how well the model performs.
Spotted a mistake or something missing on this page?Suggest an edit