Skip to main content

Evals

Evaluations

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/evals

In short

Evals are tests for AI systems: a set of inputs with expected results or grading rules, run after every change to measure how well a model or prompt performs.

What are evals?

An eval runs an AI system on a fixed set of examples and scores the results. Each example pairs an input, such as a question or a support ticket, with what a good answer looks like: an exact expected value, a list of facts it must include, or a rubric to judge it by. The score shows how often the system gets it right.

Answers from language models vary and are often free text, so evals combine several kinds of checks. Code can check structure, such as valid JSON, a correct number or a required keyword. A second model can grade an answer against a rubric, an approach called LLM-as-a-judge. People review a sample, both to catch what automated checks miss and to make sure the automated judges agree with them.

Evals play the role unit tests play in ordinary software. Teams run them before switching to a new model, editing a prompt or changing how documents are retrieved, and compare the scores to catch regressions. Production traffic is also sampled and scored, because real users ask things no test set anticipated.

Key takeaways

  • An eval scores an AI system on a fixed set of examples with known good answers or rules.
  • Checks range from exact matches and code checks to LLM-as-a-judge and human review.
  • They are run after every model, prompt or retrieval change to catch regressions.
  • A judge model must itself be checked against human ratings.

Example

A tiny eval looptypescript
const cases = [
  { input: "What is 12 + 30?", expected: "42" },
  { input: "Capital of Türkiye?", expected: "Ankara" },
];

let passed = 0;
for (const { input, expected } of cases) {
  const answer = await model.generate(input);
  if (answer.includes(expected)) passed++;
  else console.log("FAIL", input, "->", answer);
}
console.log(`${passed}/${cases.length} passed`);

Readers ask

What is LLM-as-a-judge?

It is using a language model to grade another model's answers, usually against a written rubric such as "is it correct, complete and polite?". It scales far better than human review, but the judge can be biased or wrong, so its grades should be compared with human ratings on a sample.

How are evals different from benchmarks?

Benchmarks are public, general test sets used to compare models with each other. Evals are usually your own examples, built from your product's real tasks, and they tell you whether a change makes your application better or worse.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings