GPT
Generative Pre-trained Transformer
- Pronunciation
- jee-pee-TEE
In short
GPT (Generative Pre-trained Transformer) is OpenAI's family of large language models that generate text by predicting the next token.
What is GPT?
Each word in the name describes the design. Generative: the model produces new text. Pre-trained: it first learns from a huge amount of general text before being adapted to specific uses. Transformer: it uses the transformer architecture introduced by Google researchers in 2017, whose attention mechanism lets every token look at every other token in the context.
OpenAI released GPT-1 in 2018, GPT-2 in 2019 and GPT-3 in 2020, each much larger than the last. GPT-3 showed that a big enough model could do new tasks from a few examples in the prompt. Later models were further trained on instructions and human feedback, which turned them into helpful assistants, and powered ChatGPT from November 2022 and GPT-4 in 2023.
Technically a GPT model is a decoder-only transformer. It reads the text so far as tokens and outputs a probability for every possible next token; picking one, adding it and repeating produces a reply. Developers use GPT models through OpenAI's API for chat, writing, coding, summarizing and extracting data.
A common misconception is that GPT is a synonym for every AI chatbot. GPT is OpenAI's model family; other companies build similar large language models, such as Anthropic's Claude, Google's Gemini and Meta's Llama, which share the transformer idea but are separate models.
Key takeaways
- GPT stands for Generative Pre-trained Transformer.
- It is OpenAI's family of large language models, starting with GPT-1 in 2018.
- GPT models are decoder-only transformers that predict the next token.
- Instruction tuning and human feedback turned them into assistants like ChatGPT.
- Claude, Gemini and Llama are similar LLMs, but not GPT models.
Example
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
response = client.responses.create(
model="gpt-5",
input="Explain what a REST API is in two sentences.",
)
print(response.output_text)Readers ask
What is the difference between GPT and ChatGPT?
GPT is the model: the neural network that generates text. ChatGPT is OpenAI's chat application built on top of GPT models, with a conversation interface, memory features and tools.
Is GPT the same as an LLM?
GPT models are LLMs, but not every LLM is a GPT. LLM is the general category of large language models; GPT is one family within it.
Why is it called pre-trained?
Because the model first learns general language patterns from a large body of text, and only afterwards is fine-tuned or instructed for specific tasks. The expensive general learning happens once and is reused.
See also
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- TransformerAI & Machine Learning, p. 49A transformer is a neural network architecture that uses attention to weigh how each token in a sequence relates to the others, and it powers most modern LLMs.
- ChatbotAI & Machine Learning, p. 8A chatbot is a program that converses with people in text or speech, answering questions or helping with tasks, using scripted rules or a language model.
- Generative AIAI & Machine Learning, p. 20Generative AI is artificial intelligence that creates new content, such as text, images, code, or audio, based on patterns learned from existing data.
- TokenAI & Machine Learning, p. 46A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
- Fine-tuningAI & Machine Learning, p. 19Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
Spotted a mistake or something missing on this page?Suggest an edit