Reasoning Model
- Pronunciation
- REE-zuh-ning MOD-ul
In short
A reasoning model is a language model trained to work through a problem step by step before answering, spending extra computation to do better on hard tasks.
What is a reasoning model?
An ordinary chat model starts writing its answer right away. A reasoning model first produces a hidden or summarized chain of thought: it breaks the problem down, tries approaches, checks intermediate results and corrects itself, and only then writes the reply. The idea grew from chain-of-thought prompting, where simply asking a model to think step by step improved its answers.
Reasoning models are trained for this, usually with reinforcement learning on problems whose answers can be checked, such as math problems and programming tasks with tests. OpenAI's o1 in 2024 and DeepSeek-R1 in 2025 made the approach widely known, and several model families now offer a thinking or extended reasoning mode.
The trade-off is cost and speed. Thinking uses extra tokens, so answers take longer and cost more, a pattern described as test-time compute: spending more computation while answering instead of only while training. Many APIs let developers set a thinking budget, so simple questions get quick replies and hard ones get more effort.
A common misconception is that the visible reasoning shows exactly how the model reached its answer. The written steps help, but they are generated text, not a guaranteed trace of the internal computation, and a model can still reason its way to a wrong answer. Reasoning models also bring little benefit for simple lookups or casual chat.
Key takeaways
- Reasoning models think step by step before answering.
- They are trained with reinforcement learning on checkable problems.
- OpenAI's o1 in 2024 and DeepSeek-R1 in 2025 popularized them.
- Extra thinking tokens make answers slower and more expensive.
- They help most on math, coding and multi-step problems.
Readers ask
What is the difference between a reasoning model and chain-of-thought prompting?
Chain-of-thought prompting asks any model to show its steps. A reasoning model has been specially trained to reason at length by itself, and usually does it much more reliably.
What is test-time compute?
Computation spent while the model is answering rather than while it is trained. Reasoning models use more of it by generating thinking tokens, which tends to improve results on hard problems.
When should I use a reasoning model?
For tasks with several steps where accuracy matters, such as debugging, math, planning or analysis. For quick factual answers, simple rewriting or chat, a standard model is faster and cheaper.
See also
- Chain-of-Thought PromptingAI & Machine Learning, p. 7Chain-of-thought prompting is a technique that asks an LLM to reason through intermediate steps before its final answer, improving accuracy on complex tasks.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- Reinforcement LearningAI & Machine Learning, p. 40Reinforcement learning is a type of machine learning in which an agent learns to make decisions by trial and error, earning rewards for good actions.
- InferenceAI & Machine Learning, p. 24Inference is the stage where a trained machine learning model is used to make predictions or generate output from new data, without changing what it learned.
- TokenAI & Machine Learning, p. 46A token is the basic unit of text that an LLM reads and generates, usually a whole word, part of a word, or a punctuation mark, mapped to a numeric ID.
- AI AgentAI & Machine Learning, p. 2An AI agent is a system that uses an LLM to plan and carry out multi-step tasks by deciding which tools to call, observing the results, and acting again.
Spotted a mistake or something missing on this page?Suggest an edit