Evals
Evaluations
In short
Evals are tests for AI systems: a set of inputs with expected results or grading rules, run after every change to measure how well a model or prompt performs.
What are evals?
An eval runs an AI system on a fixed set of examples and scores the results. Each example pairs an input, such as a question or a support ticket, with what a good answer looks like: an exact expected value, a list of facts it must include, or a rubric to judge it by. The score shows how often the system gets it right.
Answers from language models vary and are often free text, so evals combine several kinds of checks. Code can check structure, such as valid JSON, a correct number or a required keyword. A second model can grade an answer against a rubric, an approach called LLM-as-a-judge. People review a sample, both to catch what automated checks miss and to make sure the automated judges agree with them.
Evals play the role unit tests play in ordinary software. Teams run them before switching to a new model, editing a prompt or changing how documents are retrieved, and compare the scores to catch regressions. Production traffic is also sampled and scored, because real users ask things no test set anticipated.
Key takeaways
- An eval scores an AI system on a fixed set of examples with known good answers or rules.
- Checks range from exact matches and code checks to LLM-as-a-judge and human review.
- They are run after every model, prompt or retrieval change to catch regressions.
- A judge model must itself be checked against human ratings.
Example
const cases = [
{ input: "What is 12 + 30?", expected: "42" },
{ input: "Capital of Türkiye?", expected: "Ankara" },
];
let passed = 0;
for (const { input, expected } of cases) {
const answer = await model.generate(input);
if (answer.includes(expected)) passed++;
else console.log("FAIL", input, "->", answer);
}
console.log(`${passed}/${cases.length} passed`);Readers ask
What is LLM-as-a-judge?
It is using a language model to grade another model's answers, usually against a written rubric such as "is it correct, complete and polite?". It scales far better than human review, but the judge can be biased or wrong, so its grades should be compared with human ratings on a sample.
How are evals different from benchmarks?
Benchmarks are public, general test sets used to compare models with each other. Evals are usually your own examples, built from your product's real tasks, and they tell you whether a change makes your application better or worse.
See also
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- HallucinationAI & Machine Learning, p. 23A hallucination is when an AI model, such as an LLM, confidently produces information that sounds plausible but is false, invented, or unsupported by sources.
- Prompt EngineeringAI & Machine Learning, p. 36Prompt engineering is the practice of designing, testing, and refining the instructions given to an AI model so it produces accurate, consistent, useful output.
- Unit TestTesting & Quality, p. 35A unit test is a small, automated check that verifies one function, method, or class behaves correctly in isolation from the rest of the program.
- Regression TestingTesting & Quality, p. 22Regression testing is the practice of re-running existing tests after a code change to make sure that features which used to work have not broken.
- AI AgentAI & Machine Learning, p. 2An AI agent is a system that uses an LLM to plan and carry out multi-step tasks by deciding which tools to call, observing the results, and acting again.
- RAGAI & Machine Learning, p. 38RAG is a technique that makes an LLM answer using relevant documents retrieved at question time, so its responses are grounded in current, specific data.
- Test CaseTesting & Quality, p. 29A test case is a check of one specific behavior: starting conditions, input or steps, and the expected result that shows if the software behaves correctly.
Spotted a mistake or something missing on this page?Suggest an edit