Skip to main content

Attention Mechanism

Updated 3 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/attention-mechanism

In short

The attention mechanism is a neural network technique that lets a model decide, for each token, which other parts of the input matter most and focus on them.

What is the attention mechanism in AI?

Attention is a technique that lets a neural network look at all parts of its input and decide how much each part should influence the piece it is processing right now. It was first introduced around 2014 to improve machine translation, so a model translating a sentence could focus on the relevant source words for each word it produced. Today it is the core building block of the transformer architecture behind modern language models.

In self-attention, every token's embedding is turned into three vectors: a query, which describes what the token is looking for, a key, which describes what it offers, and a value, which carries its content. The model compares each query with every key using a dot product, converts the scores into weights that add up to 1 with the softmax function, and takes a weighted average of the values. Transformers run several of these calculations in parallel, called attention heads, so different heads can track different relationships, such as grammar, references, or topic.

A helpful analogy is a library search: your question is the query, the labels on the book spines are the keys, and the books' contents are the values. You pull the books whose labels best match your question and blend what they say into your answer. In a sentence like 'The trophy didn't fit in the suitcase because it was too big', attention is what lets the model connect 'it' with 'trophy'.

Attention is often confused with the transformer itself. Attention is one mechanism, while a transformer is a full architecture that stacks attention layers with feed-forward layers, normalization, and other parts. The name is also a metaphor, not a sign of human-like focus or understanding, and its main cost is that comparing every token with every other token grows with the square of the input length, which is why long context windows are expensive.

Key takeaways

  • Attention lets each token weigh how relevant every other token is to it.
  • It works with queries, keys, and values, combined through dot products and softmax.
  • Multi-head attention runs several attention calculations in parallel.
  • Attention is one component of a transformer, not the whole architecture.
  • Its cost grows with the square of the input length.

Example

Scaled dot-product attention in NumPypython
import numpy as np

def attention(Q, K, V):
    # Score every query against every key, scaled by the vector size
    scores = Q @ K.T / np.sqrt(K.shape[1])
    # Softmax turns each row of scores into weights that add up to 1
    weights = np.exp(scores) / np.exp(scores).sum(axis=1, keepdims=True)
    return weights @ V  # weighted average of the values

rng = np.random.default_rng(0)
tokens = rng.random((4, 8))  # 4 tokens, each an 8-number embedding
Wq, Wk, Wv = (rng.random((8, 8)) for _ in range(3))  # learned in training

output = attention(tokens @ Wq, tokens @ Wk, tokens @ Wv)
print(output.shape)  # (4, 8): one context-aware vector per token

Readers ask

What are queries, keys, and values in attention?

They are three vectors computed from each token. The query describes what a token is looking for, the key describes what a token contains, and the value is the information passed along; matching queries against keys decides how much of each value to use.

What is multi-head attention?

Multi-head attention runs several independent attention calculations, called heads, side by side, each with its own learned weights. Their results are combined, which lets the model track several kinds of relationships between tokens at once.

Why is attention expensive for long inputs?

Standard self-attention compares every token with every other token, so doubling the input length roughly quadruples the work. Models use optimized implementations, caching, and approximate forms of attention to reduce this cost.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings