Cosine Similarity
- In Turkish
- Kosinüs Benzerliği
In short
Cosine similarity measures how alike two vectors are by the angle between them, from -1 to 1; it is the usual way to compare embeddings in semantic search.
What is cosine similarity?
Cosine similarity compares the direction of two vectors and ignores their length. It is the cosine of the angle between them: 1 when they point the same way, 0 when they are at right angles and unrelated, and -1 when they point in opposite directions. It is calculated as the dot product of the vectors divided by the product of their lengths.
It matters because of embeddings. A model turns a sentence, an image or a product into a vector, and things with similar meaning end up pointing in similar directions. Comparing a question's embedding with stored document embeddings by cosine similarity finds the closest matches, which is how semantic search, recommendations and RAG retrieval usually work.
Many embedding models return vectors of length 1. For those, cosine similarity is simply the dot product, which vector databases can compute very quickly. Cosine distance, used by some tools, is 1 minus the similarity, so smaller means closer.
Key takeaways
- Cosine similarity is the cosine of the angle between two vectors, from -1 to 1.
- It compares direction and ignores length.
- It is the standard way to compare embeddings in semantic search and RAG.
- For vectors of length 1 it equals the dot product.
Example
function cosineSimilarity(a: number[], b: number[]): number {
let dot = 0, normA = 0, normB = 0;
for (let i = 0; i < a.length; i++) {
dot += a[i] * b[i];
normA += a[i] * a[i];
normB += b[i] * b[i];
}
return dot / (Math.sqrt(normA) * Math.sqrt(normB));
}
cosineSimilarity([1, 2, 3], [2, 4, 6]); // 1: same direction
cosineSimilarity([1, 0], [0, 1]); // 0: unrelatedReaders ask
Why use cosine similarity instead of the distance between points?
Because for embeddings the direction carries the meaning, while the length often reflects things such as text length. Cosine similarity ignores length, so a short and a long text about the same topic still come out as similar.
What is a good cosine similarity score?
There is no universal threshold: it depends on the embedding model and the data. Teams usually look at real examples to see where relevant and irrelevant matches separate, and pick a cut-off from that.
See also
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- Vector DatabaseAI & Machine Learning, p. 51A vector database is a database designed to store embeddings and quickly find the vectors most similar to a query, which powers semantic search and RAG.
- Semantic SearchAI & Machine Learning, p. 42Semantic search is a search technique that finds results by meaning rather than exact keywords, usually by comparing embeddings of the query and the documents.
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Natural Language ProcessingAI & Machine Learning, p. 32Natural language processing is the field of AI that teaches computers to read, understand, and generate human language in the form of text or speech.
Spotted a mistake or something missing on this page?Suggest an edit