Full-Text Search
- In Turkish
- Tam Metin Arama
In short
Full-text search is a technique that finds documents containing given words or phrases by looking them up in a text index, then ranks the results by relevance.
What is full-text search?
Full-text search lets users type words and find the records that contain them, even inside long pieces of text such as articles, product descriptions, or support tickets. Instead of looking for an exact match on a whole field, it matches individual words and ranks the results so the most relevant ones appear first. It powers the search box on most websites and apps.
It works by building an inverted index ahead of time. Each text is split into words called tokens, which are normalized by lowercasing them, removing common stop words like 'the', and reducing them to a root form through stemming, so 'running' and 'runs' both become 'run'. The index then maps every word to the list of documents that contain it, and a ranking formula such as BM25 scores each match based on how often and where the words appear.
An inverted index works like the index at the back of a textbook: rather than reading every page to find 'photosynthesis', you look up the word and jump straight to the listed pages. Many relational databases include built-in full-text search, and dedicated search engines add features such as typo tolerance, synonyms, highlighting, and faceted filters.
Full-text search is often confused with a SQL LIKE '%word%' query. LIKE looks for an exact substring, usually cannot use a normal index when the pattern starts with a wildcard, and has no idea of relevance or word forms. Full-text search is also different from semantic or vector search, which matches meaning using embeddings, so many modern systems combine both in a hybrid search.
Key takeaways
- Full-text search matches words inside text, not just exact field values.
- It relies on an inverted index that maps each word to the documents containing it.
- Tokenization, stop words, and stemming let different word forms match.
- Results are ranked by relevance, often with the BM25 formula.
- It is much faster and smarter than
LIKE '%word%'on large tables.
Example
-- Add a searchable column built from the title and body, then index it
ALTER TABLE articles
ADD COLUMN search tsvector
GENERATED ALWAYS AS (to_tsvector('english', title || ' ' || body)) STORED;
CREATE INDEX articles_search_idx ON articles USING GIN (search);
-- Find and rank articles that contain both words
SELECT title, ts_rank(search, query) AS rank
FROM articles, to_tsquery('english', 'database & index') AS query
WHERE search @@ query
ORDER BY rank DESC
LIMIT 10;Readers ask
What is the difference between full-text search and LIKE in SQL?
LIKE looks for an exact sequence of characters and usually scans the whole table when the pattern starts with %. Full-text search uses an index, understands word forms, and ranks results by relevance.
Do I need a separate search engine for full-text search?
Not always. Many databases, including PostgreSQL, MySQL, and SQLite, have built-in full-text search that is enough for many apps, while a dedicated search engine helps with very large data sets or advanced features like typo tolerance and faceting.
What is an inverted index?
An inverted index is a data structure that maps each word to the list of documents where it appears. It lets a search engine find all matching documents without scanning every text.
See also
- Database IndexDatabases, p. 7A database index is a data structure that helps a database find rows quickly without scanning a whole table, much like the index at the back of a book.
- SQLDatabases, p. 40SQL is the standard language for working with relational databases, used to create tables and to insert, query, update, and delete the data stored in them.
- DatabaseDatabases, p. 6A database is an organized collection of data stored on a computer, managed by software that lets applications save, search, and update it efficiently.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- Vector DatabaseAI & Machine Learning, p. 51A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
- ElasticsearchDatabases, p. 17Elasticsearch is a distributed search and analytics engine that indexes JSON documents for fast full-text search, filtering and aggregations over large data.
- TrieData Structures, p. 34A trie is a tree-shaped data structure that stores strings character by character, so all words that share a prefix also share the same path from the root.
Spotted a mistake or something missing on this page?Suggest an edit