Skip to main content

Transformer

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/transformer

In short

A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.

What is a transformer in AI?

A transformer is a type of neural network architecture introduced in the 2017 research paper Attention Is All You Need. It was designed for working with sequences, such as sentences, and it has become the foundation of almost all modern large language models. Transformers are also used for images, audio, source code, and even protein structures.

The key idea is self-attention. The input text is first split into tokens, and each token is turned into an embedding; then, in every layer, each token looks at all the other tokens and calculates how much attention to pay to each one. In the sentence 'The cat didn't eat because it was full', attention helps the model link the word 'it' to 'cat'. Stacking many attention layers lets the model build up a rich understanding of context.

Earlier approaches, called recurrent neural networks, read text one word at a time, which made them slow to train and prone to forgetting the beginning of long passages. Transformers look at all the tokens of the input at once, which suits parallel hardware such as GPUs and made it practical to train very large models on huge datasets. An analogy is a group discussion where everyone can hear everyone else at the same time, instead of a message passed down a line one person at a time.

A transformer is an architecture, not a product or a single model. Many different models, with different sizes and training data, are built on the same basic transformer design. Its main limitation is that the cost of attention grows quickly with input length, roughly with the square of the number of tokens, which is one reason context windows are limited.

Key takeaways

  • A transformer is a neural network architecture built around self-attention.
  • Attention lets each token weigh the relevance of every other token in the input.
  • Transformers process input tokens in parallel, which makes large-scale training practical.
  • Most modern LLMs are based on the transformer architecture.
  • Attention cost grows quickly with input length, which limits context windows.

Example

Simplified attention weights in Pythonpython
import math

def softmax(scores):
    exps = [math.exp(s) for s in scores]
    total = sum(exps)
    return [e / total for e in exps]

# Toy vectors: the query for "it" and keys for three earlier words
query_it = [1.0, 0.5]
keys = {"cat": [0.9, 0.6], "eat": [0.1, -0.4], "full": [0.4, 0.2]}

# Score = dot product of the query with each key; softmax turns scores into weights
scores = [sum(q * k for q, k in zip(query_it, key)) for key in keys.values()]
for word, weight in zip(keys, softmax(scores)):
    print(word, round(weight, 2))  # cat 0.57, eat 0.15, full 0.28

Readers ask

Why are transformers important in AI?

Transformers made it possible to train much larger language models efficiently, because they process text in parallel and capture long-range relationships between words. Nearly every modern LLM is based on this architecture.

What is self-attention?

Self-attention is the mechanism that lets each token in a sequence compute how relevant every other token is to it, and then combine their information accordingly. It is how a transformer understands context, such as which noun a pronoun refers to.

What is the difference between a transformer and an LLM?

A transformer is a neural network design, while an LLM is a specific model trained on large amounts of text. Most LLMs use the transformer architecture, but transformers are also used for images, speech, and other kinds of data.

See also

Sources

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings