Embedding
In short
An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
What is an embedding?
An embedding turns something a computer can't easily compare, such as a sentence or a photo, into a fixed-length list of numbers called a vector. An embedding model is trained so that items with similar meaning get similar vectors. For example, 'How do I reset my password?' and 'I forgot my login' produce vectors that are close together, even though they share almost no words.
You can picture embeddings as points on a map, except the map has hundreds or thousands of dimensions instead of two. To measure how related two items are, you compare their vectors, most often with cosine similarity, which looks at the angle between them. A score close to 1 means very similar meaning, while a score near 0 means the items are unrelated.
Embeddings power semantic search, recommendations, duplicate detection, clustering, and RAG systems. They are typically stored in a vector database, or in a regular database with a vector index, which can quickly find the stored vectors nearest to a query vector.
Semantic search with embeddings is different from keyword search. Keyword search matches exact words, while embedding search matches meaning, so it can find relevant results even when they use different wording. Many systems combine both approaches, which is called hybrid search.
At a glance
Key takeaways
- An embedding is a vector of numbers that captures meaning.
- Items with similar meaning have vectors that are close together.
- Cosine similarity is a common way to compare two embeddings.
- Embeddings enable semantic search, recommendations, and RAG.
- Only compare embeddings produced by the same model.
Example
import math
def cosine_similarity(a, b):
# 1.0 = same direction (similar meaning), near 0 = unrelated
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.hypot(*a) * math.hypot(*b))
# Tiny made-up embeddings (real ones have hundreds of dimensions)
cat = [0.9, 0.1, 0.3]
kitten = [0.85, 0.15, 0.35]
car = [0.1, 0.9, 0.2]
print(cosine_similarity(cat, kitten)) # about 0.996 (very similar)
print(cosine_similarity(cat, car)) # about 0.27 (not similar)Readers ask
What is a vector database?
A vector database stores embeddings and can quickly find the vectors closest to a query vector, which is called similarity or nearest-neighbor search. It is a core building block of semantic search and RAG applications.
What is the difference between an embedding and a token?
A token is a chunk of text that a language model reads, while an embedding is a vector of numbers that represents meaning. Inside an LLM each token is converted into an embedding, and dedicated embedding models produce a single vector for a whole sentence or document.
How many dimensions does an embedding have?
It depends on the model; common sizes range from a few hundred to a few thousand numbers. More dimensions can capture more nuance but need more storage and computing power.
See also
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- Database IndexDatabases, p. 7A database index is a data structure that helps a database find rows quickly without scanning a whole table, much like the index at the back of a book.
- DatabaseDatabases, p. 6A database is an organized collection of data stored on a computer, managed by software that lets applications save, search, and update it efficiently.
- Vector DatabaseAI & Machine Learning, p. 51A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
- Cosine SimilarityAI & Machine Learning, p. 13Cosine similarity measures how alike two vectors are by the angle between them, from -1 to 1; it is the usual way to compare embeddings in semantic search.
- ChunkingAI & Machine Learning, p. 9Chunking splits long documents into smaller passages before they are embedded and stored, so a RAG system can find and pass on just the relevant parts.
Sources
Spotted a mistake or something missing on this page?Suggest an edit