Open Source LangSmith Alternatives for LLM Observability

The leading open source LangSmith alternatives are Langfuse (tracing, prompt management and evals), Arize Phoenix (OpenTelemetry-based tracing and evaluation), Helicone (proxy-style logging and cost tracking) and Opik from Comet (tracing plus evaluation). All can be self-hosted and work beyond LangChain.
What is LLM observability and why do you need it?
LLM observability means recording what your AI application actually did: the prompt sent, the model and parameters, retrieved documents, tool calls, the response, latency, token usage and cost. Without it, debugging an agent is guesswork.
Traditional application monitoring tells you a request took four seconds. LLM observability tells you it took four seconds because the agent called the search tool three times with nearly identical queries and then retried a malformed JSON response.
The second job is evaluation. Once traces are stored, you can score them with rules, human review or an LLM judge, and compare prompt versions on real data instead of intuition.
Cost is the third reason. Token usage per feature, per customer and per model tells you where money goes, and which prompt changes quietly doubled your bill.
Which open source LangSmith alternatives are worth comparing?
LangSmith itself works with any framework, not only LangChain, but it is a proprietary hosted product with a self-hosted option on enterprise terms. The alternatives above give you the code and let you run it wherever your data must stay.
| Tool | How it collects data | Licence | Strongest at | Trade-off |
|---|---|---|---|---|
| Langfuse | SDKs, OpenTelemetry, framework integrations | MIT core (check the licence file for enterprise folders) | Tracing, prompt management, datasets and evals in one app | Self-hosting runs several services |
| Arize Phoenix | OpenTelemetry and OpenInference instrumentation | Elastic License 2.0 (check the licence file) | Local-first tracing and evaluation, notebook friendly | Source-available licence limits offering it as a managed service |
| Helicone | Proxy in front of the model API, or async logging | Apache 2.0 (check the licence file) | Cost, latency and request logging with minimal code changes | Proxy adds a hop in the request path |
| Opik | SDKs and framework integrations | Apache 2.0 (check the licence file) | Tracing plus evaluation metrics and experiments | Younger ecosystem than the others |
How do Langfuse and Phoenix differ?
Langfuse is a full web application for teams. You send traces from its SDKs or via OpenTelemetry, then browse sessions, inspect nested spans, manage versioned prompts, build datasets and run evaluations from the UI. It is a common default when you want one place for the whole team.
Phoenix, from Arize, grew out of a notebook-first workflow. You can launch it locally in a few lines, instrument your app with OpenInference conventions on top of OpenTelemetry, and explore traces and evaluations immediately. It also runs as a server for teams.
The licence difference matters. Langfuse’s core is MIT, while Phoenix uses the Elastic License 2.0, which allows self-hosting for your own use but restricts offering it as a hosted service to others. For internal use that rarely matters; for building an observability product, it does.
In practice many teams try both: Phoenix during early experimentation on a laptop, and Langfuse or another server tool once several people need shared access to production traces.
How do Helicone and Opik differ?
Helicone’s signature approach is the proxy. You change the base URL of your model API client to point at Helicone, add a header, and every request is logged with cost and latency. That makes it the fastest way to get visibility into an existing app.
The proxy model is weaker for deep agent traces, where you want nested spans for retrieval, tools and sub-agents. Helicone supports sessions and async logging to help, but SDK-based tools capture structure more naturally.
Opik, from Comet, focuses on tracing plus evaluation, with built-in metrics for things like hallucination and answer relevance, and experiment tracking over datasets. It suits teams that treat prompt and model changes like ML experiments.
Helicone also offers gateway-style features such as caching and rate limiting at the proxy layer, which can reduce spend on repeated requests. Evaluate those separately from logging, since they change how requests flow.
How do you choose an LLM observability tool?
- Need visibility today with minimal code? Start with a proxy approach like Helicone.
- Building multi-step agents? Prefer SDK or OpenTelemetry tracing with nested spans: Langfuse, Phoenix or Opik.
- Want prompt management in the same place? Langfuse is the strongest fit.
- Working mostly in notebooks during development? Phoenix feels natural.
- Planning to resell or host it for customers? Check licences first; permissive licences keep that door open.
- Standardise on OpenTelemetry where possible so you can switch backends later.
What does it cost to self-host LLM observability?
The licences carry no fee for self-hosting, but trace data grows fast. Every agent run can produce dozens of spans with full prompts and responses, and long-context apps make each span large.
Production-grade setups typically include an analytics database, a relational database, object storage and a queue or cache, depending on the tool. Plan storage and retention from the start: keep full payloads for a limited window and aggregated metrics longer.
Sampling helps. You rarely need every successful request in full detail; keep all errors and a representative sample of the rest.
Security is part of the cost. A trace store holds a copy of your users’ prompts and your system prompts, so it needs authentication, encryption at rest and the same access reviews as your main database.
Common mistakes with LLM tracing
- Logging secrets and personal data: prompts often contain user data. Mask sensitive fields before they reach the trace store.
- Tracing only the model call: most bugs live in retrieval and tool calls, so instrument those too.
- No session or user IDs: without them you cannot reconstruct a conversation that went wrong.
- Evals without a dataset: scoring random traces is noise. Build a curated set of real cases and grow it from failures.
- Blocking on the tracer: send traces asynchronously so an observability outage never takes down your app.
- Trusting the LLM judge blindly: spot-check its scores against human labels before using them to gate releases.
How do traces become better prompts?
The payoff comes when observability feeds a loop. A user reports a bad answer, you find the trace, add it to a dataset with the expected outcome, change the prompt or retrieval, and rerun the dataset before deploying.
Over time that dataset becomes your most valuable test suite. It encodes the real edge cases your users hit, which no generic benchmark covers.
RepoLoot’s catalog tracks observability and evaluation projects alongside agent frameworks, which makes it easier to see which tools integrate with the stack you already use.
Start small: one dataset of twenty real failures, one automated score and one human review session a week is enough to see whether a prompt change actually helped.
How do you instrument an app in an afternoon?
That last check is often the most revealing. Many teams discover duplicate model calls, unused retrieved chunks or tools called in the wrong order the first time they see a trace tree.
- Pick one backend and run it locally with the project’s Docker Compose file.
- Wrap your model client or add the SDK decorator to the main request handler.
- Add spans around retrieval, each tool call and any post-processing step.
- Attach a session ID and an anonymised user ID to every trace.
- Mask or drop fields that contain personal data before export.
- Trigger ten real requests, open the UI and check that the nested structure matches how you think the app works.
Frequently asked questions
- Does Langfuse work without LangChain?
- Yes. Langfuse has its own SDKs for Python and JavaScript, supports OpenTelemetry, and integrates with many frameworks and model providers. LangChain is just one of several integrations, so you can trace plain API calls, custom agents or other frameworks.
- Is Arize Phoenix open source?
- Phoenix’s source code is public and free to self-host, but it is released under the Elastic License 2.0, a source-available licence rather than an OSI-approved one. It allows internal use and modification but restricts offering Phoenix itself as a managed service.
- Does a proxy like Helicone add latency?
- Any proxy adds a network hop, so there is some overhead. For most applications it is small compared with model response time, and self-hosting the proxy near your app reduces it. If you want zero request-path impact, use asynchronous logging instead.
- Can I use OpenTelemetry for LLM tracing?
- Yes. OpenTelemetry has semantic conventions for generative AI, and projects like OpenInference build on it. Langfuse, Phoenix and others can ingest OpenTelemetry traces, so instrumenting with it keeps your options open if you change observability backends later.