Glossary
AI evals: how AI systems get tested
AI evals are structured tests that measure how well an AI system performs a task, usually a set of real cases with known right answers that the system is run against before release and after every change. They are how teams know whether a new prompt, model or rule made things better or worse. Without evals, an AI system's quality is a matter of opinion.
Updated · 3 min read
Kinds of evals
| Kind | What it checks | Example |
|---|---|---|
| Exact-match | Output equals the known answer | Extracted invoice total matches the ledger |
| Rule-based | Output follows a written rule | Every declined endorsement cites a reason code |
| Model-graded | Another model scores the output against a rubric | A drafted reply is accurate and on policy |
| Human review | People sample live outputs | An underwriter reviews a weekly sample of decisions |
Where the test cases come from
The best eval sets come from the business's own history: past orders, past claims, past decisions, with the outcome people actually reached. For TIE, the test set was the MGA's own historical endorsement decisions, and an agent passes when it reaches the underwriter's decision or refers the file.
Evals as a release gate
- No change to a prompt, model or rule ships without running the full set.
- Hard cases found in production are added to the set the same week.
- Results are tracked over time, so a slow decline is visible.
Sources
Frequently asked questions
What is an eval in AI?
An eval is a test of an AI system's output against a known right answer or a rule, run across many cases to measure quality.
How many test cases does an AI eval need?
Enough to cover the common cases and the known hard ones. Many teams start with a hundred or so real cases and grow the set from production.
Who should write AI evals?
Engineers build the harness, but the cases and right answers should come from the people who do the work today, because they know what correct looks like.

