Skip to main content

Cosine Similarity

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/cosine-similarity

In short

Cosine similarity measures how alike two vectors are by the angle between them, from -1 to 1; it is the usual way to compare embeddings in semantic search.

What is cosine similarity?

Cosine similarity compares the direction of two vectors and ignores their length. It is the cosine of the angle between them: 1 when they point the same way, 0 when they are at right angles and unrelated, and -1 when they point in opposite directions. It is calculated as the dot product of the vectors divided by the product of their lengths.

It matters because of embeddings. A model turns a sentence, an image or a product into a vector, and things with similar meaning end up pointing in similar directions. Comparing a question's embedding with stored document embeddings by cosine similarity finds the closest matches, which is how semantic search, recommendations and RAG retrieval usually work.

Many embedding models return vectors of length 1. For those, cosine similarity is simply the dot product, which vector databases can compute very quickly. Cosine distance, used by some tools, is 1 minus the similarity, so smaller means closer.

Key takeaways

  • Cosine similarity is the cosine of the angle between two vectors, from -1 to 1.
  • It compares direction and ignores length.
  • It is the standard way to compare embeddings in semantic search and RAG.
  • For vectors of length 1 it equals the dot product.

Example

Cosine similarity of two vectorstypescript
function cosineSimilarity(a: number[], b: number[]): number {
  let dot = 0, normA = 0, normB = 0;
  for (let i = 0; i < a.length; i++) {
    dot += a[i] * b[i];
    normA += a[i] * a[i];
    normB += b[i] * b[i];
  }
  return dot / (Math.sqrt(normA) * Math.sqrt(normB));
}

cosineSimilarity([1, 2, 3], [2, 4, 6]); // 1: same direction
cosineSimilarity([1, 0], [0, 1]);       // 0: unrelated

Readers ask

Why use cosine similarity instead of the distance between points?

Because for embeddings the direction carries the meaning, while the length often reflects things such as text length. Cosine similarity ignores length, so a short and a long text about the same topic still come out as similar.

What is a good cosine similarity score?

There is no universal threshold: it depends on the embedding model and the data. Teams usually look at real examples to see where relevant and irrelevant matches separate, and pick a cut-off from that.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings