Transformer
In short
A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
What is a transformer in AI?
A transformer is a type of neural network architecture introduced in the 2017 research paper Attention Is All You Need. It was designed for working with sequences, such as sentences, and it has become the foundation of almost all modern large language models. Transformers are also used for images, audio, source code, and even protein structures.
The key idea is self-attention. The input text is first split into tokens, and each token is turned into an embedding; then, in every layer, each token looks at all the other tokens and calculates how much attention to pay to each one. In the sentence 'The cat didn't eat because it was full', attention helps the model link the word 'it' to 'cat'. Stacking many attention layers lets the model build up a rich understanding of context.
Earlier approaches, called recurrent neural networks, read text one word at a time, which made them slow to train and prone to forgetting the beginning of long passages. Transformers look at all the tokens of the input at once, which suits parallel hardware such as GPUs and made it practical to train very large models on huge datasets. An analogy is a group discussion where everyone can hear everyone else at the same time, instead of a message passed down a line one person at a time.
A transformer is an architecture, not a product or a single model. Many different models, with different sizes and training data, are built on the same basic transformer design. Its main limitation is that the cost of attention grows quickly with input length, roughly with the square of the number of tokens, which is one reason context windows are limited.
Key takeaways
- A transformer is a neural network architecture built around self-attention.
- Attention lets each token weigh the relevance of every other token in the input.
- Transformers process input tokens in parallel, which makes large-scale training practical.
- Most modern LLMs are based on the transformer architecture.
- Attention cost grows quickly with input length, which limits context windows.
Example
import math
def softmax(scores):
exps = [math.exp(s) for s in scores]
total = sum(exps)
return [e / total for e in exps]
# Toy vectors: the query for "it" and keys for three earlier words
query_it = [1.0, 0.5]
keys = {"cat": [0.9, 0.6], "eat": [0.1, -0.4], "full": [0.4, 0.2]}
# Score = dot product of the query with each key; softmax turns scores into weights
scores = [sum(q * k for q, k in zip(query_it, key)) for key in keys.values()]
for word, weight in zip(keys, softmax(scores)):
print(word, round(weight, 2)) # cat 0.57, eat 0.15, full 0.28Readers ask
Why are transformers important in AI?
Transformers made it possible to train much larger language models efficiently, because they process text in parallel and capture long-range relationships between words. Nearly every modern LLM is based on this architecture.
What is self-attention?
Self-attention is the mechanism that lets each token in a sequence compute how relevant every other token is to it, and then combine their information accordingly. It is how a transformer understands context, such as which noun a pronoun refers to.
What is the difference between a transformer and an LLM?
A transformer is a neural network design, while an LLM is a specific model trained on large amounts of text. Most LLMs use the transformer architecture, but transformers are also used for images, speech, and other kinds of data.
See also
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- Deep LearningAI & Machine Learning, p. 14Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- TokenAI & Machine Learning, p. 46A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
- Context WindowAI & Machine Learning, p. 12A context window is the maximum amount of text, measured in tokens, that an LLM can consider at once, including the prompt, conversation history, and its reply.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- Attention MechanismAI & Machine Learning, p. 5The attention mechanism is a neural network technique that lets a model decide, for each token, which other parts of the input matter most and focus on them.
Sources
Spotted a mistake or something missing on this page?Suggest an edit