Hallucination
- In Turkish
- Halüsinasyon
In short
A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
What is an AI hallucination?
In AI, a hallucination is output that looks correct and is stated with confidence but is not true. Examples include invented facts, quotes nobody said, citations to papers that don't exist, and code that calls functions or library methods that were never defined. The term is borrowed loosely from psychology: the model is not seeing things, it is generating text that fits a familiar pattern.
Hallucinations happen because an LLM is trained to produce a likely continuation of text, not to check facts. When its training data is missing, outdated, or ambiguous on a topic, it can still write a fluent answer by filling the gaps with plausible-sounding guesses. Vague prompts, very long contexts, and questions about niche or recent topics make this more likely.
A good analogy is a student bluffing through an exam question they don't know: the answer is well written and confident, but invented. For developers, a common case is a coding assistant suggesting a package, function, or API parameter that doesn't exist, which is why generated code should always be run and tested.
A hallucination is not a bug that can be fixed in one place in the code; it is a side effect of how generative models work. Techniques such as RAG, asking the model to quote its sources, allowing it to answer that it doesn't know, and checking output with tests or tools reduce hallucinations but cannot fully eliminate them.
Key takeaways
- A hallucination is confident but false or made-up AI output.
- It happens because LLMs predict plausible text rather than verify facts.
- Invented citations, APIs, and package names are common examples.
- RAG, clear prompts, and verification reduce the risk but don't remove it.
- Always check important facts and test generated code.
Example
import json
# Suggested by an AI assistant: looks plausible, but it's a hallucination.
# Python's json module has no parse() function, so this raises AttributeError.
data = json.parse('{"name": "Ada"}')
# The real function is json.loads()
data = json.loads('{"name": "Ada"}')
print(data["name"]) # AdaReaders ask
Why do LLMs hallucinate?
LLMs generate text by predicting what is statistically likely to come next, and they have no built-in fact-checking step. When they lack reliable information about a topic, they can still produce a fluent answer that fills the gaps with invented details.
How can you reduce AI hallucinations?
Give the model relevant sources with RAG, write clear and specific prompts, ask it to cite sources or say when it doesn't know, and verify important output with tests, tools, or human review.
Can hallucinations be completely eliminated?
Not with current technology. They can be made much rarer, but any generative model can still produce incorrect output, so critical answers should always be verified.
See also
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- PromptAI & Machine Learning, p. 35A prompt is the input text or instructions you give an AI model, such as an LLM, to tell it what task to perform and what kind of answer you want.
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- EmbeddingAI & Machine Learning, p. 16An embedding is a list of numbers, called a vector, that represents the meaning of text, images, or other data so that similar items end up close together.
- EvalsAI & Machine Learning, p. 17Evals are tests for AI systems: a set of inputs with expected results or grading rules, run after every change to measure how well a model or prompt performs.
Spotted a mistake or something missing on this page?Suggest an edit