Evaluation & Testing
How to measure whether an answer is "good" when there is no single correct wording — the four-layer framework, LLM judges, and automated checks for prompts.
If you cannot measure how good your model's answers are, you cannot improve them, and you cannot safely ship them. That measuring step is called evaluation. This post builds up a full set of checks, starting from a one-line pattern match that runs in a fraction of a second and ending with one language model grading another language model against a checklist. Then we connect all of it to an automatic pipeline that stops a bad prompt before any user sees it.
In Part 1 we saw the core problem. Checking whether the output is exactly equal to a stored answer stops working the moment the output is free-form text. "Paris is the capital of France" and "The capital of France is Paris" are both correct, yet they are not the same string, so an equality check would mark one of them wrong. The solution is to run several kinds of checks, layered from cheap-and-shallow to expensive-and-deep.
Part 1 · Evaluation
The four-layer evaluation framework
A good way to organize checks is in four layers, ordered from cheapest and shallowest to most expensive and deepest. The rule is simple: run a layer only if the layers before it passed. If the output is not even valid, there is no point paying a language model to grade it. This "stop as soon as something fails" ordering keeps most checks free.
| Layer | Question it answers | Cost | Needs a known-good answer? |
|---|---|---|---|
| 1 · Structural | Is it well-formed and within the rules? | Microseconds, free | No |
| 2 · Semantic | Does it mean the same as the target? | Milliseconds, cheap | Yes |
| 3 · Model-based | Is it accurate, faithful, well-reasoned? | Seconds, costs money | Optional (a checklist) |
| 4 · Guardrail | Is it safe, fast, affordable? | Varies | No |
Before we go through the layers, one distinction that decides when each check runs. Some checks run before you ship, on a fixed set of test examples, in your build pipeline. This is called offline evaluation. Other checks run on live user traffic after you ship, watching real requests as they happen. This is called online evaluation. Layers 1 to 3 are usually run offline on test examples; Layer 4 and lighter versions of the others also run online. Part 6 covers online monitoring in detail.
Part 2 · Evaluation
Reference & semantic matching metrics
Layer 2 is used when you already have a correct answer to compare against, called the reference or gold answer. The goal is to check whether the generated text means the same thing as the reference, not whether the exact characters match. There are three families of methods, in rising order of sophistication.
2.1 Counting shared words: BLEU and ROUGE
The oldest methods just count how many short runs of words the output and the reference have in common. A run of n consecutive words is called an n-gram: "the cat" is a 2-gram, "the big cat" is a 3-gram.
The recall-focused metric, which asks "how much of the reference did the output cover?", is called ROUGE (Recall-Oriented Understudy for Gisting Evaluation). The precision-focused metric, which asks "how much of the output was actually justified by the reference?", is called BLEU (Bilingual Evaluation Understudy). Here is a tiny worked example.
Reference: the cat sat on the mat (6 words)
Output: the cat sat on a rug (6 words)
Shared words: the, cat, sat, on = 4 matches.
ROUGE would report the recall side (0.67), BLEU the precision side (0.67). These are fast and need no model, but they only see shared words. A perfect paraphrase that reuses none of the same words, like "a rug is where the animal rested", would score close to zero even though the meaning is fine. So these metrics work for summarization and translation, where wording is expected to stay close, and work poorly for open-ended questions.
2.2 Comparing meaning with embeddings
A better approach turns each piece of text into a list of numbers that captures its meaning. That list of numbers is called an embedding, and the model that produces it is an embedding model. Two texts that mean similar things get embeddings that point in a similar direction. We measure how similar two directions are with the cosine similarity: the cosine of the angle between the two lists of numbers. It runs from 1 (same direction, same meaning) down toward 0 (unrelated).
Here \(\mathbf v_g\) is the generated answer's embedding and \(\mathbf v_r\) is the reference's. The top is the dot product (multiply matching numbers, add them up). The bottom divides by each list's length so only direction matters, not size.
Real embeddings have 1,000 or more numbers; we use 3 here so the arithmetic is visible. The math is identical. Reference embedding \(\mathbf v_r=[0.9,0.3,0.1]\).
$$\begin{aligned} \text{Output A (a paraphrase) } [0.8,0.4,0.1]:\ &\cos=\tfrac{(0.9)(0.8)+(0.3)(0.4)+(0.1)(0.1)}{\sqrt{0.91}\,\sqrt{0.81}}=\tfrac{0.85}{0.858}\approx \mathbf{0.99}\\[4pt] \text{Output B (off-topic) } [0.1,0.2,0.95]:\ &\cos=\tfrac{(0.9)(0.1)+(0.3)(0.2)+(0.1)(0.95)}{\sqrt{0.91}\,\sqrt{0.9525}}=\tfrac{0.245}{0.931}\approx \mathbf{0.26} \end{aligned}$$The paraphrase scores 0.99 even though it shares few exact words, and the off-topic answer scores 0.26. This is exactly what word-counting could not do. In practice you pick a cutoff, for example "accept if cosine is above 0.82", tuned on your own data.
# minimal embedding-similarity check
from sentence_transformers import SentenceTransformer, util
model = SentenceTransformer("all-MiniLM-L6-v2")
gen = "The capital of France is Paris."
ref = "Paris is the capital of France."
v_g, v_r = model.encode([gen, ref])
score = util.cos_sim(v_g, v_r).item() # ~0.94
assert score > 0.82 # passes: same meaning, different word order
2.3 BERTScore — matching word by word, in context
The previous method squeezes a whole sentence into one embedding, which can blur details. A refinement embeds every word in its context, then matches each output word to the reference word it is most similar to, and averages those similarities. This method is called BERTScore (it uses a model called BERT, Bidirectional Encoder Representations from Transformers). Because it compares word by word in context, it lines up with human judgment better than BLEU or ROUGE on many tasks.
Output a fast dog vs reference a quick dog. Match each output word to its closest reference word and read off the cosine similarity:
| Output word | Best reference match | Cosine |
|---|---|---|
| a | a | 1.00 |
| fast | quick | 0.87 |
| dog | dog | 1.00 |
"fast" never appears in the reference, so plain word-counting scores it zero; BERTScore sees that "fast" and "quick" are close in meaning and gives 0.87. Recall is computed the same way but matching each reference word to its closest output word, and F1 combines the two.
Part 3 · Evaluation
LLM-as-a-judge — power and pitfalls
The deepest and most flexible check is to hand the output to another language model and ask it to grade the answer against a set of instructions. This is called LLM-as-a-judge (LLM stands for Large Language Model). It can rate things no formula can, such as accuracy, tone, helpfulness, or whether the answer stuck to a source. It is also the easiest to trust too much, because these judges have consistent, predictable biases that you have to actively cancel out.
3.1 The two ways to ask a judge
- Give one answer a score: "Rate this answer from 1 to 5 for factual accuracy." Simple, but the scores tend to drift over time and bunch up in the middle. This is called absolute scoring.
- Compare two answers: "Here are answers A and B. Which is better?" More reliable, because deciding which of two things is better is a steadier judgment than putting a number on one thing alone. This is called pairwise comparison, and it is the basis of the A/B regression testing in Part 6.
Here is a concrete judge prompt for the pairwise case.
SYSTEM: You are a strict grader. You will see a QUESTION, a
REFERENCE answer, and two candidate answers A and B. Decide
which candidate better matches the reference on factual accuracy.
Ignore length and writing style. Do not reward extra detail that
is not in the reference. Reply with exactly one token: A, B, or TIE.
QUESTION: What is the capital of France?
REFERENCE: Paris.
A: The capital of France is Paris.
B: France is a country in Western Europe with many large cities.
ASSISTANT: A
3.2 G-Eval — a more reliable recipe
A widely used improvement on plain judging is called G-Eval. It adds two ideas. First, before grading, it asks the judge to write out its own step-by-step checklist from the criteria, so the grading is more consistent. Second, instead of taking the single score the model happened to output, it looks at how sure the model was about each possible score and takes a weighted average. The model's confidence in each score is called its probability for that score, and turning raw model outputs into probabilities that add up to 1 is done with the softmax function.
Suppose the judge is 70% sure the answer is a 4 and 30% sure it is a 5, and sure of nothing else. A plain judge would just say "4". G-Eval computes:
$$\text{score}=(1)(0)+(2)(0)+(3)(0)+(4)(0.70)+(5)(0.30)=2.8+1.5=\mathbf{4.3}$$The 4.3 records that the judge leaned slightly toward 5, information the flat "4" would have thrown away. Over a whole test set, these finer scores make small quality changes visible.
3.3 The three biases you must correct
A "bias" here means a consistent way the judge is unfair that has nothing to do with actual answer quality. There are three you should assume are present.
| Bias | What you see | How to correct it |
|---|---|---|
| Position bias | The judge tends to prefer whichever answer it was shown first. | Run the comparison both ways (A then B, then B then A) and average. Or only count a win if the answer wins in both orders. |
| Verbosity bias | Longer answers get higher scores even when they are not better. | Tell the judge in the prompt to ignore length, and dock points for padding in the checklist. |
| Self-preference bias | A judge tends to favor answers written by its own model family. | Use a different model as the judge than the one being tested, and average several different judges. |
3.4 RAG-specific model-based metrics
Many systems answer a question by first fetching relevant documents and then writing an answer using them. This is called RAG (Retrieval-Augmented Generation), and Part 5 covers it fully. The fetched documents are called the retrieved context. Three model-based checks are the standard way to grade a RAG system; together they are often called the "RAG triad" and are implemented by tools named Ragas and TruLens. Each works by having a language model break the text into small factual statements, called claims, and checking each one.
Answer: "Aspirin reduces fever and cures diabetes." Break into claims:
- Claim 1: "Aspirin reduces fever." — supported by the retrieved context. ✓
- Claim 2: "Aspirin cures diabetes." — not in the context (made up). ✗
A score of 0.5 flags that half the answer was invented, so it should not ship.
- Faithfulness (also called groundedness): the fraction of the answer's claims that the retrieved documents actually support. This is the main signal for catching made-up content, called hallucination.
- Answer relevancy: does the answer actually address the question? Ragas estimates this by asking a model to invent questions the answer would fit, then measuring the cosine similarity between those invented questions and the real one.
- Context precision and recall: did the retrieval step fetch the right documents? Precision asks whether the fetched documents were relevant and well-ranked; recall asks whether it fetched all the relevant ones. These separate a retrieval problem from a writing problem.
Part 4 · Evaluation
Golden datasets — the asset that compounds
None of the metrics above mean anything without a set of examples to run them on. That curated set of test cases — each one an input paired with its expected output and the way it should be graded — is called a golden dataset. It is the most valuable and longest-lived asset in an LLMOps project. Models change, prompts change, vendors change; the golden dataset is the fixed ruler you measure all of them against.
4.1 What goes into it
- Common requests (happy paths): the everyday, typical inputs. Usually the smallest and least interesting part.
- Edge cases: very long inputs, empty inputs, vague requests, other languages, and inputs written to trip the system up.
- Regression cases: every past production failure, frozen forever as a permanent test. A "regression" is when a change re-breaks something that used to work; this is the feedback loop from Part 1 made concrete.
- A grading rule per case: exact match, a keyword that must appear, a cosine-similarity cutoff, or a judge checklist. The grading method is stored alongside each example.
Part 5 · Evaluation
Testing the LLM layers of an agentic app
Evaluating a plain model asks "is the text good?" But many real applications do more than write text. An agent is a system that uses a model to decide which actions to take, calling tools, querying APIs or MCP servers (MCP means Model Context Protocol, a standard way to expose tools to a model), and carrying out multi-step plans. Testing an agent asks more questions: did it pick the right tool, pass valid inputs to it, take a sensible sequence of steps, and follow the rules of its domain?
The standard way to organize these tests is a testing pyramid: many fast cheap tests at the bottom, fewer slow expensive tests at the top.
5.1 Tier 1 — plain unit tests
Because the output follows a fixed format (from Part 1), the tool inputs and the final output are typed, so an ordinary assertion checks them with no model call at all.
def test_listing_structure_and_constraints():
out = run_listing_agent("Men's running shoe, blue, lightweight, size 10")
assert len(out.title) <= 80 # marketplace hard limit
assert len(out.bullet_points) >= 3 # minimum feature bullets
assert out.category in ALLOWED_TAXONOMY
5.2 Tier 2 — path and tool tests
The sequence of steps an agent takes is called its trajectory. An agent can reach a right-looking answer by a wrong path. To test this cheaply and repeatably, you replace the real tools with stand-ins that return fixed responses, called mocks, then check the exact list of calls the agent made.
def test_mcp_tool_trajectory():
tracer = AgentTracer()
run_listing_agent("Men's running shoe", tracer=tracer)
# Correct grounding: look up the taxonomy BEFORE writing copy
assert "mcp_get_category_taxonomy" in tracer.tool_calls
assert tracer.index("mcp_get_category_taxonomy") < tracer.index("generate_copy")
assert tracer.tool_call_count < 5 # no runaway tool loop
This catches two failures that an output-only test would miss: the agent skipping a required lookup step and then guessing the category, and the agent getting stuck in a loop that burns time and money.
5.3 Tier 3 — domain evaluation, with two worked case studies
Case A · E-commerce product-listing agent
The top-tier test checks that the agent did not invent product features that were not in the supplier's text (a hallucination).
def test_no_hallucinated_attributes():
score = evaluate_hallucination(
input="Lightweight mesh canvas shoe",
output=generated_listing,
context="Mesh canvas is not waterproof.")
assert score < 0.1 # agent must not claim 'waterproof'
Case B · Medical literature-synthesis agent
Here the stakes make Layers 1 and 2 not enough on their own; the ways it can fail are specific to the domain and unsafe.
| Failure mode | Why it is dangerous | How to test it |
|---|---|---|
| Fake citations | Invents believable-looking paper IDs that do not exist. | Direct verification: read out each citation, call the PubMed API, and assert the paper exists and matches the claim. |
| Flipped negation | Turns "not indicated" into "indicated" — invisible to cosine similarity. | Contradiction check: a classifier is asked "Does output X contradict established fact Y?" This is called natural language inference, or NLI. |
| Wrong terminology | Uses casual words instead of the official medical terms. | Dictionary check: verify key terms against a controlled vocabulary such as MeSH. |
Part 6 · Evaluation
CI/CD for prompts & agents
All of the above only helps if it runs automatically on every change. Running your tests automatically whenever code changes is called CI/CD (Continuous Integration and Continuous Delivery). The pattern is the same as ordinary software CI/CD, with two extra twists for language models: you must control cost, and you compare the new version's scores against the current production version rather than just checking pass/fail.
6.1 The tool landscape
Here is a minimal example of one of these frameworks, DeepEval, which lets you write evals in the same style as ordinary unit tests.
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import GEval
correctness = GEval(
name="Correctness",
criteria="Is the answer factually correct vs the reference?",
threshold=0.7, # fail the test below this score
)
def test_capital_question():
case = LLMTestCase(
input="What is the capital of France?",
actual_output="The capital of France is Paris.",
expected_output="Paris.")
assert_test(case, [correctness]) # runs the judge, asserts score >= 0.7
| Tool | Best for | Notes |
|---|---|---|
| Promptfoo | Testing many prompts and models from a config file; attack testing | Declarative YAML; easy to run in CI. |
| DeepEval | Unit-test-style evals, G-Eval, RAG and agent metrics | Feels like writing unit tests; wide set of metrics. |
| Ragas | RAG faithfulness, relevancy, and context metrics | The reference implementation for RAG evaluation. |
| OpenAI Evals | A registry of reusable evals and model-graded templates | Open-source framework plus a shared registry. |
| Braintrust / LangSmith | Hosted evaluation, dataset management, and logging | Connect offline evals to live monitoring (Part 6). |
Appendix
References & further reading
Grouped by type; freely available online. Tool capabilities and metric names evolve — verify specifics against current official docs.
Research papers & primary sources
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. https://arxiv.org/abs/2303.16634
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Position/verbosity/self-preference biases. https://arxiv.org/abs/2306.05685
- BERTScore: Evaluating Text Generation with BERT. https://arxiv.org/abs/1904.09675
- RAGAS: Automated Evaluation of Retrieval Augmented Generation. https://arxiv.org/abs/2309.15217
- BLEU: a Method for Automatic Evaluation of Machine Translation. https://aclanthology.org/P02-1040/
Tools & documentation
- DeepEval documentation. https://docs.confident-ai.com/
- Ragas documentation. https://docs.ragas.io/
- Test your prompts, models, and RAGs. https://www.promptfoo.dev/
- The RAG Triad. https://www.trulens.org/
Named tools are illustrative of the state of the field as of early 2026, not endorsements or fixed facts. The four-layer framework and testing pyramid are conceptual syntheses of common industry practice.