Semantic Search
- In Turkish
- Anlamsal Arama
In short
Semantic search is a search technique that finds results by meaning rather than exact keywords, usually by comparing embeddings of the query and the documents.
What is semantic search?
Semantic search returns results that match what a query means, not just the words it contains. A search for 'how to cancel my plan' can find an article titled 'Ending your subscription', even though the two share no important words. It works by representing both queries and documents as embeddings, lists of numbers that capture meaning, and looking for the documents whose embeddings are closest to the query's.
A typical setup has two stages. Ahead of time, documents are split into chunks, each chunk is turned into an embedding by an embedding model, and the vectors are stored in a vector database or a vector index inside a regular database. At query time, the query is embedded with the same model, the nearest vectors are found, often with cosine similarity, and the top results may be reordered by a slower but more precise model called a reranker.
The difference is like looking something up in the index at the back of a book versus asking a knowledgeable librarian. The index only helps if you know the exact word the author used, while the librarian understands what you're after and points you to the right chapter. Semantic search powers site and help-center search, product and code search, duplicate detection, recommendations, and the retrieval step of RAG systems.
Semantic search is often contrasted with full-text search. Full-text search matches keywords using an inverted index and is excellent for exact terms like product codes, error messages, and names, while semantic search handles paraphrases and natural questions better but can miss exact identifiers, so many systems combine both in hybrid search. Semantic search is also not the same as a vector database: the database is a storage and indexing tool, while semantic search is the technique that uses it.
Key takeaways
- Semantic search matches by meaning, so different wording can still find the right result.
- Queries and documents are compared as embeddings, usually with cosine similarity.
- Always embed queries and documents with the same model.
- Keyword search is better for exact terms; hybrid search combines both.
- It is the usual retrieval step in RAG systems.
Example
# embed() and cosine_similarity() are placeholders for an embedding model and a vector helper
docs = [
"Ending your subscription",
"Changing your profile picture",
"Refund policy for annual plans",
]
doc_vectors = [embed(d) for d in docs] # computed once and stored
query_vector = embed("how to cancel my plan")
scores = [cosine_similarity(query_vector, v) for v in doc_vectors]
# Rank documents by meaning, not by shared keywords
ranked = sorted(zip(scores, docs), reverse=True)
print(ranked[0][1]) # Ending your subscriptionReaders ask
What is the difference between semantic search and keyword search?
Keyword search finds documents that contain the query's words, while semantic search finds documents with a similar meaning, even if they use different words. Keyword search is better for exact names and codes; semantic search is better for natural-language questions.
What is hybrid search?
Hybrid search runs keyword search and semantic search together and merges their results into one ranking. It combines the precision of exact matching with the flexibility of matching by meaning.
Do I need a vector database for semantic search?
Not always. Small collections can be searched by comparing vectors directly in memory, and many regular databases support vector columns and indexes. A dedicated vector database helps mainly with very large collections or strict latency requirements.
See also
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- Vector DatabaseAI & Machine Learning, p. 51A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
- Full-Text SearchDatabases, p. 22Full-text search is a technique that finds documents containing given words or phrases by looking them up in a text index, then ranks the results by relevance.
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- Natural Language ProcessingAI & Machine Learning, p. 32Natural language processing is the field of AI that teaches computers to read, understand, and generate human language in the form of text or speech.
- Database IndexDatabases, p. 7A database index is a data structure that helps a database find rows quickly without scanning a whole table, much like the index at the back of a book.
- Cosine SimilarityAI & Machine Learning, p. 13Cosine similarity measures how alike two vectors are by the angle between them, from -1 to 1; it is the usual way to compare embeddings in semantic search.
- ChunkingAI & Machine Learning, p. 9Chunking splits long documents into smaller passages before they are embedded and stored, so a RAG system can find and pass on just the relevant parts.
Spotted a mistake or something missing on this page?Suggest an edit