Elasticsearch
- Pronunciation
- ee-LAS-tik-surch
In short
Elasticsearch is a distributed search and analytics engine that indexes JSON documents for fast full-text search, filtering and aggregations over large data.
What is Elasticsearch?
Elasticsearch was first released in 2010 by Shay Banon and is built on Apache Lucene, a search library. You send it JSON documents, such as products, articles or log lines, and it builds an inverted index: a map from every word to the documents that contain it. That index is what lets it find matching documents among millions in milliseconds.
Searches go beyond exact matches. Elasticsearch splits text into words, can ignore case and word endings, ranks results by relevance, tolerates typos and highlights the matching parts. It also aggregates, for example counting orders per country or showing error rates per minute, which is why it is used for analytics as well as search.
Data is split into shards spread across the nodes of a cluster, with replica copies for safety, so it scales by adding machines. It is often used as part of the Elastic Stack, formerly called ELK: Logstash or Beats collect logs, Elasticsearch stores and indexes them, and Kibana shows them in dashboards.
A common misconception is that Elasticsearch should be the main database. It is usually a secondary store, filled from the main database or a log pipeline, because it trades some consistency for search speed. Its license changed in 2021, which led AWS to start OpenSearch, a fork that works in much the same way.
Key takeaways
- Elasticsearch is a distributed search and analytics engine built on Lucene.
- An inverted index maps words to documents for fast full-text search.
- It ranks by relevance, tolerates typos and computes aggregations.
- Data is split into shards across a cluster and replicated.
- It usually complements a main database instead of replacing it.
Example
# Add a document to the "products" index
curl -X POST "localhost:9200/products/_doc" -H "Content-Type: application/json" -d '
{ "name": "Wireless keyboard", "price": 49 }'
# Full-text search; "keybord" still matches thanks to fuzziness
curl -X GET "localhost:9200/products/_search" -H "Content-Type: application/json" -d '
{ "query": { "match": { "name": { "query": "keybord", "fuzziness": "AUTO" } } } }'Readers ask
Is Elasticsearch a database?
It stores and retrieves data, so it can be called a database, but it is designed as a search and analytics engine and is usually kept alongside a primary database that remains the source of truth.
What is the ELK stack?
ELK stands for Elasticsearch, Logstash and Kibana: Logstash collects and processes logs, Elasticsearch indexes them and Kibana visualizes them. With Beats added, it is now called the Elastic Stack.
What is OpenSearch?
OpenSearch is an open-source fork of Elasticsearch started by AWS in 2021, after Elastic changed its license. It works in a very similar way and is now maintained under the Linux Foundation.
See also
- Full-Text SearchDatabases, p. 22Full-text search is a technique that finds documents containing given words or phrases by looking them up in a text index, then ranks the results by relevance.
- Document DatabaseDatabases, p. 15A document database is a NoSQL database that stores each record as a self-contained document, usually JSON-like, whose fields can differ from record to record.
- ShardingDatabases, p. 39Sharding is a way of scaling a database by splitting its data across several servers, called shards, so each one stores and handles only part of the total.
- LoggingDevOps & Cloud, p. 35Logging is the practice of recording timestamped messages about events in a running program, such as errors and requests, so people can investigate them later.
- ObservabilityDevOps & Cloud, p. 38Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- JSONBackend & APIs, p. 25JSON is a lightweight, text-based format for storing and exchanging structured data as key-value pairs and lists, readable by both humans and machines.
Spotted a mistake or something missing on this page?Suggest an edit