Insights / How-To Guides
How to Test an AI Feature Before Your Customers Do: A Practical Guide to LLM Evals
AI features fail quietly. This guide shows how to build an evaluation set, grade outputs automatically and with humans, and catch regressions before every release.
By Syntax Station Engineering · · 3 min read
Key takeaways
- An eval is a fixed set of real inputs with known good outcomes that you run on every change.
- Combine three graders: exact checks in code, a model acting as judge, and periodic human review.
- Grade retrieval and generation separately so you know which part to fix.
- Production logs are your best source of new test cases.
Traditional software fails loudly: an error, a crash, a red test. AI features fail quietly. The answer is fluent, formatted and wrong, and nobody notices until a customer does.
Evaluations, or "evals", are how serious teams catch these failures before release. They are not complicated, but they need to exist before you start tuning prompts.
What an eval set looks like
An eval set is a list of cases. Each case has:
- An input: a user question, a document, a ticket.
- Context, if relevant: the user's account type, the documents available.
- An expected outcome or grading rule: the correct answer, required facts, a forbidden action, or a rubric.
Cases should come from reality. Use past tickets, real documents and actual user questions, not examples someone invented on a whiteboard.
Three kinds of graders
Code checks
Anything you can check in code, check in code. Is the output valid JSON? Does it include the order number? Did the agent call the refund tool when it should not have? Is the response under the length limit? These checks are fast, cheap and unambiguous.
Model-graded checks
For open-ended answers, a second model can grade the output against a rubric: "Does the answer state the correct notice period? Does it cite a source? Is the tone appropriate?" Keep rubrics specific. "Is this a good answer?" produces noisy grades. "Does the answer mention the 30-day limit?" does not.
Human review
People remain the final authority. Review a sample of outputs regularly, especially for new features and sensitive topics, and use those reviews to calibrate your model graders.
Grade each stage separately
For a retrieval-based assistant, a wrong answer can come from two places: the right document was never found, or it was found and the model ignored or misread it. Measure both:
| Stage | Example metric |
|---|---|
| Retrieval | Was the correct passage in the top results? |
| Generation | Is every claim in the answer supported by the passages? |
| Behavior | Did it decline when it should have? Did it hand off correctly? |
| Cost and speed | Tokens, cost and latency per case |
Run evals on every change
Changes that seem harmless can shift behavior: a new model version, an edited prompt, a different chunk size, an updated document. Run the full eval set in your CI pipeline and compare against the last release. Block the release if key scores drop.
Feed production back into the set
After launch, logs become your richest source of test cases. Every time a user flags a bad answer or a reviewer spots one, add it to the eval set with the correct outcome. Over time the set reflects how people really use the feature.
Common mistakes
- Testing only happy paths. Include ambiguous questions, off-topic requests, angry users and attempts to make the assistant misbehave (see our guide to prompt injection).
- Changing the test set and the system at the same time. You will not know which change moved the score.
- Chasing a single number. An overall score can hide a serious regression in one category. Report by category.
Where to start this week
Take 50 real inputs, write down what a good outcome looks like for each, and run them through your current system. The first run is usually humbling and always useful. It turns "the AI seems fine" into a number your team can improve.
Frequently asked questions
What is an LLM eval?
An evaluation (eval) is a repeatable test of an AI system: a set of inputs, the expected outcomes or grading rules, and a score. It works like a unit test suite for behavior that is not fully deterministic.
How many test cases do I need?
Start with 50 to 100 well-chosen cases covering common requests and known edge cases. Grow it to several hundred as you collect failures from production.
Can I use an AI model to grade another AI model?
Yes, and it scales well, but calibrate it. Compare the judge model's grades with human grades on a sample and refine its rubric until they agree closely.