LLM
Large Language Model
In short
An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
What is an LLM?
A large language model, or LLM, is a neural network trained on a very large collection of text, such as books, websites, and source code. It learns the patterns of language well enough to answer questions, summarize documents, translate, and write code. The word large refers both to the amount of training data and to the number of parameters, which can reach billions or more.
Under the hood, an LLM works with tokens, which are small chunks of text such as words or parts of words. Given some input text, called a prompt, the model predicts a likely next token, adds it to the text, and repeats the process until the answer is complete. Most modern LLMs are based on the transformer architecture, which lets the model weigh how much each earlier token matters when predicting the next one.
Developers usually use an LLM through an API: they send a prompt and receive generated text in return. LLMs power chat assistants, coding assistants, customer support bots, and search features. A useful mental model is a very well-read autocomplete: it is excellent at producing plausible text, but it does not look facts up unless it is connected to a source of information.
An LLM is not a database or a search engine. It does not reliably store documents word for word, its knowledge stops at a training cutoff date, and it can produce confident but wrong answers, known as hallucinations. Techniques like RAG are used to ground its answers in real, up-to-date data.
At a glance
Key takeaways
- An LLM generates text by predicting the next token over and over.
- It is trained on massive text datasets and has billions of parameters.
- Its built-in knowledge is frozen at a training cutoff date.
- Output quality depends heavily on the prompt and context you provide.
- LLMs can hallucinate, so important answers should be verified.
Example
// Send a prompt to an LLM through a generic HTTP API
const response = await fetch("https://llm.example.com/v1/generate", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
prompt: "Explain recursion in one sentence.",
maxTokens: 60, // limit the length of the answer
temperature: 0.2, // lower = more predictable output
}),
});
const data = await response.json();
console.log(data.text);Readers ask
What is the difference between an LLM and AI?
AI is the broad field of making machines perform intelligent tasks; an LLM is one specific kind of AI model focused on understanding and generating text. Chat assistants are applications built on top of LLMs.
What is a token in an LLM?
A token is the unit of text an LLM reads and writes, often a whole word or part of a word. In English, one token is roughly three quarters of a word on average, and usage limits and pricing are usually measured in tokens.
What is a context window?
The context window is the maximum amount of text, measured in tokens, that an LLM can take into account at once, including both the prompt and its answer. Anything outside the window is invisible to the model.
See also
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- PromptAI & Machine Learning, p. 35A prompt is the input text or instructions you give an AI model, such as an LLM, to tell it what task to perform and what kind of answer you want.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- HallucinationAI & Machine Learning, p. 23A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
- APIBackend & APIs, p. 2An API is a set of rules that lets one piece of software request data or actions from another in a predictable, documented way.
- GPTAI & Machine Learning, p. 21GPT (Generative Pre-trained Transformer) is OpenAI's family of large language models that generate text by predicting the next token.
- Reasoning ModelAI & Machine Learning, p. 39A reasoning model is a language model trained to work through a problem step by step before answering, spending extra computation to do better on hard tasks.
- Mixture of ExpertsAI & Machine Learning, p. 28A mixture of experts (MoE) is a neural network design that sends each input to only a few of many small experts, so a huge model costs far less to run.
- EvalsAI & Machine Learning, p. 17Evals are tests for AI systems: a set of inputs with expected results or grading rules, run after every change to measure how well a model or prompt performs.
- Attention MechanismAI & Machine Learning, p. 5The attention mechanism is a neural network technique that lets a model decide, for each token, which other parts of the input matter most and focus on them.
Sources
Spotted a mistake or something missing on this page?Suggest an edit