Best open source LLM evaluation tools

7 minUpdated:
Best open source LLM evaluation tools

Use promptfoo for prompt and regression testing in CI, DeepEval for pytest-style assertions on LLM apps, Ragas for measuring RAG retrieval and faithfulness, and lm-evaluation-harness for benchmarking models themselves. Most teams combine an app-level tool with a small, hand-labelled test set.

What does an LLM evaluation tool actually do?

An evaluation tool runs a set of inputs through your prompt, chain, agent or model, then scores the outputs. Scores come from exact checks, heuristics, embedding similarity or another model acting as a judge.

The point is to replace “it looked fine when I tried it” with a repeatable number you can compare between versions. Without that, every prompt tweak or model swap is a guess.

There are two different jobs hiding under the same word. Model evaluation compares base models on standard benchmarks. Application evaluation checks whether your specific product behaves correctly on your data.

Most teams need application evaluation first. Model benchmarks matter only when you are choosing or training the model itself.

Which open source evaluation tools should you know?

ToolLanguageEvaluatesStyleBest forTrade-off
promptfooTypeScript, CLIPrompts, apps, providersYAML test configs, CLI, web viewerRegression tests and red-teaming in CIConfig files grow large on complex apps
DeepEvalPythonLLM apps, RAG, agentspytest-like test cases and metricsPython teams that already use pytestMany metrics rely on an LLM judge, which costs tokens
RagasPythonRAG pipelinesMetrics over question, context, answerRetrieval quality and faithfulnessNarrower scope than general tools
lm-evaluation-harnessPythonModelsStandard academic benchmarksComparing base or fine-tuned modelsSays little about your specific app
OpenAI EvalsPythonModels and promptsRegistry of eval definitionsReusing community eval patternsOriented around OpenAI-style APIs
InspectPythonModels and agentsTasks, solvers, scorersStructured agent and safety evalsMore framework concepts to learn

When is promptfoo the right choice?

promptfoo treats prompts like code under test. You declare prompts, providers and test cases with assertions, run a command, and get a matrix of results you can view in the browser or fail a CI job on.

It shines when you compare several models or prompt variants side by side, and it includes red-teaming features for probing jailbreaks and prompt injection. It is language-agnostic because it can call any HTTP endpoint.

Because configs are plain files, they live in the same repository as your prompts and go through code review. That makes prompt changes auditable in the same way as code changes.

When is DeepEval or Ragas better?

DeepEval fits Python codebases. You write test cases in code, attach metrics such as answer relevancy, faithfulness or custom judge criteria, and run them with a pytest-style runner.

Ragas focuses on retrieval-augmented generation. Its metrics separate retrieval problems from generation problems: did the retriever find relevant context, and did the answer stay faithful to it? That split tells you whether to fix chunking or the prompt.

Both lean on model-graded metrics. Pin the judge model and its version, or your scores will drift for reasons unrelated to your change.

Do you need lm-evaluation-harness?

Only if you train, fine-tune or select base models. lm-evaluation-harness runs a large set of standard benchmarks against local or hosted models with consistent settings, which is how many public leaderboards are produced.

For a product team calling an API, benchmark scores are a weak signal. A model that ranks higher on general reasoning can still do worse on your support tickets or your extraction schema.

It is still worth knowing, because it explains how public model scores are produced and why small changes in prompt format or few-shot settings can move them noticeably. That context helps you read leaderboards with appropriate scepticism.

How to set up LLM evaluation step by step

  • Collect 30 to 100 real inputs from logs or users, including awkward edge cases and past failures.
  • Write the expected behaviour for each: an exact answer, required facts, a JSON schema, or a rubric.
  • Start with cheap deterministic checks such as schema validity, required keywords and length limits.
  • Add model-graded metrics only where rules cannot judge quality, and spot-check the judge against human labels.
  • Run the suite on every prompt, model or retrieval change, and block merges on regressions.
  • Feed new production failures back into the test set every week.

Where LLM evaluation breaks

  • Test sets written by the developer who wrote the prompt, which only test what they already expected.
  • Trusting a single aggregate score that hides a collapse in one category of questions.
  • Using the same model as generator and judge, which tends to favour its own style.
  • Letting the test set go stale while real traffic shifts to new topics.
  • Ignoring cost and latency, which are part of quality for users.

How these tools fit with observability

Evaluation runs before release; observability watches production. Tools such as Langfuse or Phoenix capture traces from live traffic, and those traces become the raw material for your next test cases.

A healthy loop looks like this: trace in production, label failures, add them to the eval suite, fix, and confirm the score moved. The evaluation tool is only as useful as the data you keep feeding it.

Which metrics matter for different LLM apps?

The right metric depends on what the app does. Generic “quality” scores hide the failure modes that users actually notice, so choose checks that map to your task.

App typeUseful checksTypical tool
Extraction to JSONSchema validity, field-level exact matchpromptfoo or DeepEval assertions
RAG question answeringContext recall, faithfulness, answer relevancyRagas or DeepEval RAG metrics
Support assistantPolicy compliance, tone rubric, escalation correctnessLLM judge with a written rubric
Agents with toolsCorrect tool chosen, valid arguments, task completedInspect or custom traces plus assertions
Model selectionBenchmark accuracy under fixed settingslm-evaluation-harness

How do you evaluate agents rather than single prompts?

Agents fail in the middle, not only at the end. An agent can reach a correct final answer after calling the wrong tool three times, which costs money and time and may have side effects.

Evaluate the trajectory as well as the result. Record every tool call with its arguments, then assert on the ones that matter: was the refund tool called only after the lookup, were arguments valid, did the agent stop when it should.

Run agent evals in a sandbox with fake tools or test accounts. Evaluating against production systems turns a test suite into an incident.

How much does running evaluations cost?

The tools are free, but model calls are not. Each run costs one generation per test case plus one judge call per model-graded metric. That grows quickly when you test several prompt variants against several models.

Keep costs down by running the cheap deterministic suite on every commit and the full judge-based suite nightly or before releases. Cache outputs for unchanged prompts, and use a smaller judge for coarse checks. Track eval spend like any other CI expense.

Who should own the evaluation suite?

Evaluation works best when it has an owner and the people who know the domain help write it. Engineers can build the harness, but support leads, lawyers or analysts know which answers are actually wrong.

Give domain experts a simple way to label outputs, even a spreadsheet, and turn their labels into test cases. Review the suite in the same rhythm as product changes, and retire tests that no longer reflect how the product is used.

Treat a failing eval like a failing unit test: someone investigates, the fix is recorded, and the test stays in the suite so the problem cannot quietly return.

Frequently asked questions

What is the best open source LLM evaluation tool?
There is no single best. promptfoo is the easiest way to add regression tests to CI in any language, DeepEval suits Python teams, Ragas is specialised for RAG, and lm-evaluation-harness is the standard for comparing models. Many teams run two of them together.
Is LLM-as-a-judge reliable?
It is useful but imperfect. Judges can prefer longer answers, their own style, or the first option shown. Use clear rubrics, pin the judge model, and compare a sample of judge scores with human labels before trusting the numbers for release decisions.
How many test cases do I need to evaluate an LLM app?
Start with a few dozen well-chosen real examples that cover your main tasks and known failures. That is enough to catch obvious regressions. Grow the set over time from production failures rather than generating hundreds of synthetic cases upfront.
Can I run LLM evaluations without paying for API calls?
Yes, partly. Deterministic checks cost nothing, and you can point judge metrics at a local model served by Ollama or vLLM. Local judges are usually weaker than frontier models, so validate them against human labels before relying on their scores.
Free for builders

Get a hand-picked shortlist of repos for your project

Tell us what you are building. A person — not a bot — reviews it and replies within 48 hours with the catalog projects that fit, including licence and difficulty notes.

We use your email only for this request. Privacy policy

Related guides