How to Evaluate an LLM Application Before Production

6 minUpdated:
How to Evaluate an LLM Application Before Production

Build a test set from real or realistic inputs with expected outcomes, score outputs with code checks first and LLM-as-judge only where needed, calibrate judges against human labels, run the suite on every prompt or model change, and keep sampling production traffic after launch.

What does it mean to evaluate an LLM application?

Evaluating an LLM application is not the same as benchmarking a model. Public benchmarks tell you how a model does on someone else’s tasks. Your evaluation tells you whether your prompts, retrieval, tools and model together do your task well enough to ship.

The output of a good evaluation is a decision: ship, do not ship, or ship with limits. If your eval results never change a decision, they are decoration.

Think of it as a test suite that tolerates fuzziness. Some checks are exact, such as valid JSON or a correct order ID. Others are graded, such as whether an answer is faithful to the source documents.

How do you build an evaluation dataset?

Start from real inputs whenever you can: support tickets, search logs, beta user sessions or internal dogfooding. Synthetic questions generated by a model are useful for coverage, but they tend to be cleaner and easier than what users actually type.

Label each case with what a good outcome looks like. That might be a reference answer, a list of facts that must appear, a tool that must be called, or simply “must refuse.” Keep the labels short so reviewers can agree on them.

Cover the edges on purpose. Include ambiguous questions, out-of-scope requests, adversarial inputs, long conversations and inputs in every language you support. A hundred well-chosen cases beat a thousand near-duplicates.

  • Happy path: the common questions that make up most traffic.
  • Hard but valid: multi-step, long-context or rare-but-important tasks.
  • Should refuse or hand off: out-of-scope, unsafe or unanswerable inputs.
  • Adversarial: prompt injection attempts and attempts to extract the system prompt.
  • Regression cases: every bug a user has reported, added the day it was fixed.

Which metrics should you use?

Choose metrics that map to a user-visible failure. For a retrieval-based assistant that usually means: did retrieval find the right documents, is the answer faithful to them, and does it actually answer the question. For an agent it means: did it complete the task, with the right tools, within budget.

Always track operational metrics next to quality: latency, token usage, cost per request and error rate. A prompt change that improves answers but doubles latency is a trade-off someone must decide on, not a free win.

Metric typeHow to score itGood forWatch out for
Exact and structural checksCode: schema validation, regex, equalityFormats, IDs, tool choice, refusalsOnly covers what you can specify precisely
Reference similarityCompare with a gold answerShort factual answersPenalizes correct answers phrased differently
Faithfulness or groundednessJudge model checks claims against sourcesRAG assistantsJudge must see the same sources
Task successCode or judge checks the end stateAgents and workflowsNeeds a realistic sandbox or mocks
Human preferenceSide-by-side ratings by reviewersTone, helpfulness, styleSlow and costly; use for calibration
OperationalLogs and tracesLatency, cost, errorsEasy to ignore until it is a problem

When is LLM-as-judge reliable?

Using a model to grade outputs scales well, but a judge is itself an LLM application that needs evaluating. Judges are known to prefer longer answers, to be swayed by confident tone, and in pairwise comparisons to favor a particular position.

Make judges narrow. Ask one yes-or-no question per call, such as “Does the answer contain any claim not supported by the sources?”, and require a short justification. Broad “rate this answer from 1 to 10” prompts produce noisy scores.

Calibrate before trusting. Have humans label a sample, run the judge on the same sample, and look at where they disagree. Re-check whenever you change the judge prompt or the judge model.

Which open-source tools help?

You can run evaluations with a spreadsheet and a script, and many teams start that way. Tools help once you want repeatable runs, comparisons between versions and a place for reviewers to label outputs.

ToolWhat it doesBest forTrade-off
promptfooConfig-driven test cases and assertions across prompts and modelsCI checks for promptsLess suited to deep trace analysis
RagasMetrics aimed at retrieval-augmented generationRAG pipelinesJudge-based metrics need calibration
DeepEvalPytest-style tests with LLM metricsPython teams that like unit testsMetric defaults may not fit your task
LangfuseTracing, datasets, scores and human annotationLinking evals to production tracesA service to self-host or pay for
InspectFramework for structured model and agent evaluationsAgent tasks and safety-style evalsMore setup than a simple test runner

How do you run evaluations step by step?

  • Write down the three to five failures that would make the product unacceptable, and make sure each has test cases.
  • Assemble the first dataset, even if it is only fifty cases, and store it in version control next to the prompts.
  • Implement code-based checks first, then add judge-based checks only for qualities code cannot measure.
  • Record a baseline with the current prompt and model before changing anything.
  • Run the suite on every change to prompts, retrieval settings, tools or model version, and compare against the baseline.
  • Set release gates, for example no regression on must-refuse cases and no drop in task success beyond an agreed tolerance.
  • Run the suite several times on important changes, because model outputs vary between runs.
  • After launch, sample production traffic, label it, and feed failures back into the dataset.

Where does it break? Common mistakes

  • Evaluating only on examples the team wrote while building the prompt, which measures memory rather than quality.
  • Averaging everything into one score, hiding a collapse in a small but critical category.
  • Treating a single run as the truth when outputs are non-deterministic.
  • Changing the dataset and the prompt at the same time, so you cannot tell which caused the score change.
  • Letting a judge model grade outputs from the same model family without any human calibration.
  • Stopping evaluation at launch, even though model updates, new documents and new users keep changing behavior.

How do you evaluate retrieval and agents separately?

End-to-end scores tell you something went wrong, not where. Split the pipeline and score each stage. For retrieval, check whether the documents that contain the answer appear in the top results; you only need a list of relevant document IDs per test question to do this in code.

For generation, hold retrieval fixed by feeding the model the correct documents, then check whether it answers faithfully. If the answer is good with perfect context but bad in the full pipeline, the problem is retrieval, not the prompt.

Agents need trajectory checks as well as outcome checks. Record which tools were called in what order, and flag runs that reached the right answer through a dangerous or wasteful path, such as calling a write tool when a read tool was enough.

Mock external systems during these tests. A sandbox with fixed responses makes runs repeatable and prevents an evaluation from sending real email or creating real records.

How do you monitor quality after launch?

Production is the only place where you see the real distribution of inputs. Log inputs, outputs, retrieved context and tool calls with enough metadata to reproduce a case, while respecting your privacy commitments.

Combine cheap automatic signals with human review. Thumbs-down clicks, repeated rephrasing, handoff requests and abandoned sessions point you to cases worth reading. Weekly review of a random sample catches failures users never report.

Pin model versions where your provider allows it, and treat a provider’s model update like a dependency upgrade: run the suite before switching.

Frequently asked questions

How many test cases do I need to evaluate an LLM app?
Enough to cover every important category with several cases each. Many teams start with fifty to a few hundred cases and grow from production failures. Coverage of distinct behaviors matters more than raw count, and a small set you actually run on every change beats a large one nobody maintains.
Can I use the same model as both generator and judge?
You can, but it adds risk that the judge shares the generator’s blind spots or favors its style. Prefer narrow judge questions, a different model where practical, and calibration against human labels. For high-stakes checks, keep humans reviewing a sample regardless of which judge you use.
Should evaluations run in CI?
Yes for fast, cheap checks such as formatting, tool choice and a core regression set. Larger judge-based suites can run nightly or before releases to control cost and time. Store results so you can compare versions and see when a regression was first introduced.
What is the difference between offline and online evaluation?
Offline evaluation runs a fixed dataset before release, so you can compare versions under identical conditions. Online evaluation measures real traffic after release through user signals, sampling and review. You need both: offline to catch regressions early, online to discover the failures your dataset did not anticipate.
Free for builders

Get a hand-picked shortlist of repos for your project

Tell us what you are building. A person — not a bot — reviews it and replies within 48 hours with the catalog projects that fit, including licence and difficulty notes.

We use your email only for this request. Privacy policy

Related guides