LLMOps, In Depth · Part 2 of 6

Evaluation & Testing

How to measure whether an answer is "good" when there is no single correct wording — the four-layer framework, LLM judges, and automated checks for prompts.

By PrithvirajPart 2 of 6~34 min read

If you cannot measure how good your model's answers are, you cannot improve them, and you cannot safely ship them. That measuring step is called evaluation. This post builds up a full set of checks, starting from a one-line pattern match that runs in a fraction of a second and ending with one language model grading another language model against a checklist. Then we connect all of it to an automatic pipeline that stops a bad prompt before any user sees it.

In Part 1 we saw the core problem. Checking whether the output is exactly equal to a stored answer stops working the moment the output is free-form text. "Paris is the capital of France" and "The capital of France is Paris" are both correct, yet they are not the same string, so an equality check would mark one of them wrong. The solution is to run several kinds of checks, layered from cheap-and-shallow to expensive-and-deep.

The core idea
A traditional software test asks a yes/no question: is the output equal to the expected answer? An LLM evaluation asks a different, richer question: how good is the output, and good in what specific ways? The answer is not one true/false value but several graded scores, for example one score for accuracy, one for tone, one for whether it stuck to the source. Everything below is a way to compute those scores cheaply, reliably, and in a way that cannot be easily fooled.

Part 1 · Evaluation

The four-layer evaluation framework

A good way to organize checks is in four layers, ordered from cheapest and shallowest to most expensive and deepest. The rule is simple: run a layer only if the layers before it passed. If the output is not even valid, there is no point paying a language model to grade it. This "stop as soon as something fails" ordering keeps most checks free.

Layer 1 · Structural checks
Is the shape right? Valid JSON, correct length, allowed values. Runs in microseconds, no model needed.
↓ if valid
Layer 2 · Semantic checks
Does it mean the same as a known-good answer? Word overlap, embedding similarity. Cheap, needs a reference answer.
↓ if it matches
Layer 3 · Model-based checks
Is it accurate and well-reasoned? A language model grades against a checklist. Slow and costs money.
↓ always, in parallel
Layer 4 · Guardrail checks
Is it safe, fast, affordable? Private data, toxicity, attacks, latency, cost. Covered in Part 6.
LayerQuestion it answersCostNeeds a known-good answer?
1 · StructuralIs it well-formed and within the rules?Microseconds, freeNo
2 · SemanticDoes it mean the same as the target?Milliseconds, cheapYes
3 · Model-basedIs it accurate, faithful, well-reasoned?Seconds, costs moneyOptional (a checklist)
4 · GuardrailIs it safe, fast, affordable?VariesNo

Before we go through the layers, one distinction that decides when each check runs. Some checks run before you ship, on a fixed set of test examples, in your build pipeline. This is called offline evaluation. Other checks run on live user traffic after you ship, watching real requests as they happen. This is called online evaluation. Layers 1 to 3 are usually run offline on test examples; Layer 4 and lighter versions of the others also run online. Part 6 covers online monitoring in detail.

In practice — most failures are caught at Layer 1
It is tempting to jump straight to having a language model grade everything, but that is the slowest and most expensive tool you have. If you designed your output to follow a fixed format (covered in Part 1), then Layer 1 by itself catches cut-off answers, broken JSON, out-of-range numbers, and missing fields. Those are a large share of real-world breakages, and they cost nothing to catch. Save your expensive model-grading budget for the answers that already passed the cheap layers.

Part 2 · Evaluation

Reference & semantic matching metrics

Layer 2 is used when you already have a correct answer to compare against, called the reference or gold answer. The goal is to check whether the generated text means the same thing as the reference, not whether the exact characters match. There are three families of methods, in rising order of sophistication.

2.1 Counting shared words: BLEU and ROUGE

The oldest methods just count how many short runs of words the output and the reference have in common. A run of n consecutive words is called an n-gram: "the cat" is a 2-gram, "the big cat" is a 3-gram.

The recall-focused metric, which asks "how much of the reference did the output cover?", is called ROUGE (Recall-Oriented Understudy for Gisting Evaluation). The precision-focused metric, which asks "how much of the output was actually justified by the reference?", is called BLEU (Bilingual Evaluation Understudy). Here is a tiny worked example.

Worked example — counting shared 1-grams

Reference: the cat sat on the mat (6 words)
Output: the cat sat on a rug (6 words)

Shared words: the, cat, sat, on = 4 matches.

$$\text{precision}=\frac{4\ \text{matched}}{6\ \text{output words}}=0.67 \qquad \text{recall}=\frac{4\ \text{matched}}{6\ \text{reference words}}=0.67$$

ROUGE would report the recall side (0.67), BLEU the precision side (0.67). These are fast and need no model, but they only see shared words. A perfect paraphrase that reuses none of the same words, like "a rug is where the animal rested", would score close to zero even though the meaning is fine. So these metrics work for summarization and translation, where wording is expected to stay close, and work poorly for open-ended questions.

2.2 Comparing meaning with embeddings

A better approach turns each piece of text into a list of numbers that captures its meaning. That list of numbers is called an embedding, and the model that produces it is an embedding model. Two texts that mean similar things get embeddings that point in a similar direction. We measure how similar two directions are with the cosine similarity: the cosine of the angle between the two lists of numbers. It runs from 1 (same direction, same meaning) down toward 0 (unrelated).

$$\text{cosine similarity}=\cos\theta=\frac{\mathbf v_g\cdot\mathbf v_r}{\lVert\mathbf v_g\rVert\,\lVert\mathbf v_r\rVert}$$

Here \(\mathbf v_g\) is the generated answer's embedding and \(\mathbf v_r\) is the reference's. The top is the dot product (multiply matching numbers, add them up). The bottom divides by each list's length so only direction matters, not size.

Worked example — paraphrase scores high, wrong answer scores low

Real embeddings have 1,000 or more numbers; we use 3 here so the arithmetic is visible. The math is identical. Reference embedding \(\mathbf v_r=[0.9,0.3,0.1]\).

$$\begin{aligned} \text{Output A (a paraphrase) } [0.8,0.4,0.1]:\ &\cos=\tfrac{(0.9)(0.8)+(0.3)(0.4)+(0.1)(0.1)}{\sqrt{0.91}\,\sqrt{0.81}}=\tfrac{0.85}{0.858}\approx \mathbf{0.99}\\[4pt] \text{Output B (off-topic) } [0.1,0.2,0.95]:\ &\cos=\tfrac{(0.9)(0.1)+(0.3)(0.2)+(0.1)(0.95)}{\sqrt{0.91}\,\sqrt{0.9525}}=\tfrac{0.245}{0.931}\approx \mathbf{0.26} \end{aligned}$$

The paraphrase scores 0.99 even though it shares few exact words, and the off-topic answer scores 0.26. This is exactly what word-counting could not do. In practice you pick a cutoff, for example "accept if cosine is above 0.82", tuned on your own data.

# minimal embedding-similarity check
from sentence_transformers import SentenceTransformer, util

model = SentenceTransformer("all-MiniLM-L6-v2")
gen  = "The capital of France is Paris."
ref  = "Paris is the capital of France."

v_g, v_r = model.encode([gen, ref])
score = util.cos_sim(v_g, v_r).item()   # ~0.94
assert score > 0.82   # passes: same meaning, different word order

2.3 BERTScore — matching word by word, in context

The previous method squeezes a whole sentence into one embedding, which can blur details. A refinement embeds every word in its context, then matches each output word to the reference word it is most similar to, and averages those similarities. This method is called BERTScore (it uses a model called BERT, Bidirectional Encoder Representations from Transformers). Because it compares word by word in context, it lines up with human judgment better than BLEU or ROUGE on many tasks.

Worked example — how BERTScore matches words

Output a fast dog vs reference a quick dog. Match each output word to its closest reference word and read off the cosine similarity:

Output wordBest reference matchCosine
aa1.00
fastquick0.87
dogdog1.00
$$\text{precision}=\frac{1.00+0.87+1.00}{3}=0.96$$

"fast" never appears in the reference, so plain word-counting scores it zero; BERTScore sees that "fast" and "quick" are close in meaning and gives 0.87. Recall is computed the same way but matching each reference word to its closest output word, and F1 combines the two.

Engineering track — what meaning-based checks cannot catch
Embedding similarity measures whether two texts are about the same thing, not whether they are true. Consider "The drug is indicated for hypertension" and "The drug is not indicated for hypertension". They share almost every word and point in nearly the same direction, so cosine similarity is high, yet they say opposite things. A single flipped word like "not", a wrong number, or a subtle reversal of a fact all pass Layer 2 unnoticed. That blind spot is the whole reason Layer 3 exists, and it is why a safety-critical domain like the medical example in Part 5 cannot stop at meaning-based similarity.

Part 3 · Evaluation

LLM-as-a-judge — power and pitfalls

The deepest and most flexible check is to hand the output to another language model and ask it to grade the answer against a set of instructions. This is called LLM-as-a-judge (LLM stands for Large Language Model). It can rate things no formula can, such as accuracy, tone, helpfulness, or whether the answer stuck to a source. It is also the easiest to trust too much, because these judges have consistent, predictable biases that you have to actively cancel out.

3.1 The two ways to ask a judge

  • Give one answer a score: "Rate this answer from 1 to 5 for factual accuracy." Simple, but the scores tend to drift over time and bunch up in the middle. This is called absolute scoring.
  • Compare two answers: "Here are answers A and B. Which is better?" More reliable, because deciding which of two things is better is a steadier judgment than putting a number on one thing alone. This is called pairwise comparison, and it is the basis of the A/B regression testing in Part 6.

Here is a concrete judge prompt for the pairwise case.

SYSTEM: You are a strict grader. You will see a QUESTION, a
REFERENCE answer, and two candidate answers A and B. Decide
which candidate better matches the reference on factual accuracy.
Ignore length and writing style. Do not reward extra detail that
is not in the reference. Reply with exactly one token: A, B, or TIE.

QUESTION: What is the capital of France?
REFERENCE: Paris.
A: The capital of France is Paris.
B: France is a country in Western Europe with many large cities.

ASSISTANT: A

3.2 G-Eval — a more reliable recipe

A widely used improvement on plain judging is called G-Eval. It adds two ideas. First, before grading, it asks the judge to write out its own step-by-step checklist from the criteria, so the grading is more consistent. Second, instead of taking the single score the model happened to output, it looks at how sure the model was about each possible score and takes a weighted average. The model's confidence in each score is called its probability for that score, and turning raw model outputs into probabilities that add up to 1 is done with the softmax function.

$$\text{score}=\sum_{i=1}^{5} i\cdot p(i),\qquad p(i)=\frac{e^{z_i}}{\sum_{j=1}^{5} e^{z_j}}$$
Worked example — why the weighted score is finer

Suppose the judge is 70% sure the answer is a 4 and 30% sure it is a 5, and sure of nothing else. A plain judge would just say "4". G-Eval computes:

$$\text{score}=(1)(0)+(2)(0)+(3)(0)+(4)(0.70)+(5)(0.30)=2.8+1.5=\mathbf{4.3}$$

The 4.3 records that the judge leaned slightly toward 5, information the flat "4" would have thrown away. Over a whole test set, these finer scores make small quality changes visible.

3.3 The three biases you must correct

A "bias" here means a consistent way the judge is unfair that has nothing to do with actual answer quality. There are three you should assume are present.

BiasWhat you seeHow to correct it
Position biasThe judge tends to prefer whichever answer it was shown first.Run the comparison both ways (A then B, then B then A) and average. Or only count a win if the answer wins in both orders.
Verbosity biasLonger answers get higher scores even when they are not better.Tell the judge in the prompt to ignore length, and dock points for padding in the checklist.
Self-preference biasA judge tends to favor answers written by its own model family.Use a different model as the judge than the one being tested, and average several different judges.

3.4 RAG-specific model-based metrics

Many systems answer a question by first fetching relevant documents and then writing an answer using them. This is called RAG (Retrieval-Augmented Generation), and Part 5 covers it fully. The fetched documents are called the retrieved context. Three model-based checks are the standard way to grade a RAG system; together they are often called the "RAG triad" and are implemented by tools named Ragas and TruLens. Each works by having a language model break the text into small factual statements, called claims, and checking each one.

$$\text{Faithfulness}=\frac{\text{number of claims in the answer that the retrieved context supports}}{\text{total number of claims in the answer}}$$
Worked example — computing faithfulness

Answer: "Aspirin reduces fever and cures diabetes." Break into claims:

  • Claim 1: "Aspirin reduces fever." — supported by the retrieved context. ✓
  • Claim 2: "Aspirin cures diabetes." — not in the context (made up). ✗
$$\text{Faithfulness}=\frac{1\ \text{supported}}{2\ \text{total}}=0.5$$

A score of 0.5 flags that half the answer was invented, so it should not ship.

  • Faithfulness (also called groundedness): the fraction of the answer's claims that the retrieved documents actually support. This is the main signal for catching made-up content, called hallucination.
  • Answer relevancy: does the answer actually address the question? Ragas estimates this by asking a model to invent questions the answer would fit, then measuring the cosine similarity between those invented questions and the real one.
  • Context precision and recall: did the retrieval step fetch the right documents? Precision asks whether the fetched documents were relevant and well-ranked; recall asks whether it fetched all the relevant ones. These separate a retrieval problem from a writing problem.

Part 4 · Evaluation

Golden datasets — the asset that compounds

None of the metrics above mean anything without a set of examples to run them on. That curated set of test cases — each one an input paired with its expected output and the way it should be graded — is called a golden dataset. It is the most valuable and longest-lived asset in an LLMOps project. Models change, prompts change, vendors change; the golden dataset is the fixed ruler you measure all of them against.

4.1 What goes into it

  • Common requests (happy paths): the everyday, typical inputs. Usually the smallest and least interesting part.
  • Edge cases: very long inputs, empty inputs, vague requests, other languages, and inputs written to trip the system up.
  • Regression cases: every past production failure, frozen forever as a permanent test. A "regression" is when a change re-breaks something that used to work; this is the feedback loop from Part 1 made concrete.
  • A grading rule per case: exact match, a keyword that must appear, a cosine-similarity cutoff, or a judge checklist. The grading method is stored alongside each example.
Intuition — 100 great cases beat 10,000 sloppy ones
The same rule that governs training data (Part 3) governs test data: quality matters more than quantity. A tight set of 100 to 300 carefully labeled, varied, and deliberately tricky cases gives a sharper, faster, and cheaper signal than tens of thousands of scraped examples with fuzzy expected answers. Start with 50 hand-written cases and grow the set from real failures, not from bulk-generated filler.
In practice — the golden dataset lives in version control
Store it in your code repository right next to the prompts (Part 1). When you change a prompt, the change in eval scores is your evidence to either merge or block the change. When a new model version comes out, re-running the golden dataset is your migration test. When a customer reports a bug, the reproduction becomes case number 247 in the set, and no future change can silently break it again.

Part 5 · Evaluation

Testing the LLM layers of an agentic app

Evaluating a plain model asks "is the text good?" But many real applications do more than write text. An agent is a system that uses a model to decide which actions to take, calling tools, querying APIs or MCP servers (MCP means Model Context Protocol, a standard way to expose tools to a model), and carrying out multi-step plans. Testing an agent asks more questions: did it pick the right tool, pass valid inputs to it, take a sensible sequence of steps, and follow the rules of its domain?

The standard way to organize these tests is a testing pyramid: many fast cheap tests at the bottom, fewer slow expensive tests at the top.

Tier 3 · few, slow
Whole-task domain evaluation
Did it accomplish the task and follow domain rules? Graded by an LLM judge plus rule checks, on the golden dataset.
Tier 2 · some
Path (trajectory) and tool-choice tests
Did it call the right tool, in the right order, with valid inputs, without looping forever? Fake tool responses so the test is deterministic.
Tier 1 · many, fast
Unit tests on shape and inputs
Plain assertions on the structured output and the tool-call inputs. No model call.

5.1 Tier 1 — plain unit tests

Because the output follows a fixed format (from Part 1), the tool inputs and the final output are typed, so an ordinary assertion checks them with no model call at all.

def test_listing_structure_and_constraints():
    out = run_listing_agent("Men's running shoe, blue, lightweight, size 10")
    assert len(out.title) <= 80          # marketplace hard limit
    assert len(out.bullet_points) >= 3    # minimum feature bullets
    assert out.category in ALLOWED_TAXONOMY

5.2 Tier 2 — path and tool tests

The sequence of steps an agent takes is called its trajectory. An agent can reach a right-looking answer by a wrong path. To test this cheaply and repeatably, you replace the real tools with stand-ins that return fixed responses, called mocks, then check the exact list of calls the agent made.

def test_mcp_tool_trajectory():
    tracer = AgentTracer()
    run_listing_agent("Men's running shoe", tracer=tracer)
    # Correct grounding: look up the taxonomy BEFORE writing copy
    assert "mcp_get_category_taxonomy" in tracer.tool_calls
    assert tracer.index("mcp_get_category_taxonomy") < tracer.index("generate_copy")
    assert tracer.tool_call_count < 5      # no runaway tool loop

This catches two failures that an output-only test would miss: the agent skipping a required lookup step and then guessing the category, and the agent getting stuck in a loop that burns time and money.

5.3 Tier 3 — domain evaluation, with two worked case studies

Case A · E-commerce product-listing agent

The top-tier test checks that the agent did not invent product features that were not in the supplier's text (a hallucination).

def test_no_hallucinated_attributes():
    score = evaluate_hallucination(
        input="Lightweight mesh canvas shoe",
        output=generated_listing,
        context="Mesh canvas is not waterproof.")
    assert score < 0.1   # agent must not claim 'waterproof'
Case B · Medical literature-synthesis agent

Here the stakes make Layers 1 and 2 not enough on their own; the ways it can fail are specific to the domain and unsafe.

Failure modeWhy it is dangerousHow to test it
Fake citationsInvents believable-looking paper IDs that do not exist.Direct verification: read out each citation, call the PubMed API, and assert the paper exists and matches the claim.
Flipped negationTurns "not indicated" into "indicated" — invisible to cosine similarity.Contradiction check: a classifier is asked "Does output X contradict established fact Y?" This is called natural language inference, or NLI.
Wrong terminologyUses casual words instead of the official medical terms.Dictionary check: verify key terms against a controlled vocabulary such as MeSH.
In practice — use the cheapest layer that catches the failure
Notice that the medical tests are not all model-grading calls. A fake citation is caught by a plain API lookup (Layer 1 style). Wrong terminology is caught by a dictionary. Only the flipped-negation case actually needs a model. Good agent testing means matching each way the system can fail to the shallowest, cheapest check that reliably catches it, and saving expensive model-grading for the failures nothing cheaper can find.

Part 6 · Evaluation

CI/CD for prompts & agents

All of the above only helps if it runs automatically on every change. Running your tests automatically whenever code changes is called CI/CD (Continuous Integration and Continuous Delivery). The pattern is the same as ordinary software CI/CD, with two extra twists for language models: you must control cost, and you compare the new version's scores against the current production version rather than just checking pass/fail.

Someone proposes a prompt or agent change (a pull request)
↓
1 · Fast structural tests
Check the shape of the output. Runs on every case in seconds, no model.
↓
2 · Mock-tool path tests
Tool choice and order, with a small cheap model checking the inputs.
↓
3 · Golden dataset run + judge
Meaning checks plus LLM-as-a-judge on the 50 to 300 saved cases.
↓
4 · Compare scores against the live version
New prompt vs current production, on the exact same cases.
↓ scores high enough and nothing regressed?
✓ Merge  |  ✗ Block the change and attach a report

6.1 The tool landscape

Here is a minimal example of one of these frameworks, DeepEval, which lets you write evals in the same style as ordinary unit tests.

from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import GEval

correctness = GEval(
    name="Correctness",
    criteria="Is the answer factually correct vs the reference?",
    threshold=0.7,   # fail the test below this score
)

def test_capital_question():
    case = LLMTestCase(
        input="What is the capital of France?",
        actual_output="The capital of France is Paris.",
        expected_output="Paris.")
    assert_test(case, [correctness])   # runs the judge, asserts score >= 0.7
ToolBest forNotes
PromptfooTesting many prompts and models from a config file; attack testingDeclarative YAML; easy to run in CI.
DeepEvalUnit-test-style evals, G-Eval, RAG and agent metricsFeels like writing unit tests; wide set of metrics.
RagasRAG faithfulness, relevancy, and context metricsThe reference implementation for RAG evaluation.
OpenAI EvalsA registry of reusable evals and model-graded templatesOpen-source framework plus a shared registry.
Braintrust / LangSmithHosted evaluation, dataset management, and loggingConnect offline evals to live monitoring (Part 6).

Appendix

References & further reading

Grouped by type; freely available online. Tool capabilities and metric names evolve — verify specifics against current official docs.

Research papers & primary sources

  1. Liu et al. (2023). G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. https://arxiv.org/abs/2303.16634
  2. Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Position/verbosity/self-preference biases. https://arxiv.org/abs/2306.05685
  3. Zhang et al. (2019). BERTScore: Evaluating Text Generation with BERT. https://arxiv.org/abs/1904.09675
  4. Es et al. (2023). RAGAS: Automated Evaluation of Retrieval Augmented Generation. https://arxiv.org/abs/2309.15217
  5. Papineni et al. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation. https://aclanthology.org/P02-1040/

Tools & documentation

  1. Confident AI. DeepEval documentation. https://docs.confident-ai.com/
  2. Exploding Gradients. Ragas documentation. https://docs.ragas.io/
  3. Promptfoo. Test your prompts, models, and RAGs. https://www.promptfoo.dev/
  4. TruLens. The RAG Triad. https://www.trulens.org/

Named tools are illustrative of the state of the field as of early 2026, not endorsements or fixed facts. The four-layer framework and testing pyramid are conceptual syntheses of common industry practice.