Back to all posts
Guide
8 min read

Best LLM Observability Tools 2026: Langfuse, Helicone, and LangSmith Compared

DevToolLab Team

DevToolLab Team

June 29, 2026

Best LLM Observability Tools 2026: Langfuse, Helicone, and LangSmith Compared

A production LLM app making 50,000 API calls a day can silently hallucinate on half of them. Datadog tells you the HTTP status was 200 - it has no idea the model confidently made something up. Standard APM tools were built for request latency and error rates. They have three specific blind spots for LLM workloads:

  • Token costs are not request costs. A single request can cost $0.001 or $0.40 depending on input length. You can be within rate limits and 10x over budget.
  • Errors are soft. A hallucination returns HTTP 200 with a plausible-looking body. Detecting it requires evaluation, not exception handling.
  • Multi-step agents are invisible. An agent that makes 5 tool calls produces 6 model calls. Without trace stitching, you see 6 independent requests and no picture of what the agent actually did.

In 2026, this has a dedicated category: LLM observability - covering tracing, evals, cost attribution, and prompt management. This guide covers the five most widely used tools with code, real pricing, and honest tradeoffs.

Quick Comparison

ToolArchitectureSelf-hostedBest ForFree TierStarting Price
LangfuseSDK + OTELYesFull tracing, prompt mgmt, evalsYes (50k obs/mo)$59/mo
HeliconeProxyYesDrop-in logging, cost trackingYes (100k req/mo)$20/mo
BraintrustSDKNoEvaluation-first, prompt testingYes (5 users)$200/mo
Phoenix by ArizeOTEL-nativeYesLocal debugging, open sourceFree (open source)Free/$99/mo
LangSmithSDKPartialLangChain/LangGraph appsYes (5k runs/mo)$39/mo

Langfuse

Langfuse
Langfuse

Langfuse is the most widely adopted open-source LLM observability platform. It covers traces, prompt management, evaluations, datasets, and user session analytics - and the entire codebase is MIT-licensed, so you can self-host with no data leaving your environment.

The core abstraction is the trace: a tree of observations that maps to one user interaction. An agent making 4 LLM calls to answer one question produces one trace with 4 generation spans. Each span captures the exact prompt sent, the response, token counts, latency, and model name.

Python
from langfuse.openai import openai  # drop-in wrapper

client = openai.OpenAI()

response = client.chat.completions.create(
    model="gpt-5.4",
    messages=[{"role": "user", "content": "Explain async/await in Python"}],
    metadata={"feature": "code-explainer", "user_id": "usr_123"},
)

Prompt management is one of Langfuse's strongest features - store prompts in Langfuse and pull the current version at runtime. Rolling out a prompt change becomes a dashboard operation, not a deploy. Self-hosting runs on Docker Compose (git clone, docker compose up -d, done).

Pricing: Hobby free (50k observations/month) - Pro $59/month - Team $499/month. Self-hosted is always free.

What it does well: The most complete feature set in the category. Traces, prompt versioning, datasets, human-in-the-loop evals, LLM-as-judge pipelines, and cost attribution in one tool. Self-hosting is well documented and production-tested.

What it does not do: Real-time alerting is limited on the free tier. Self-hosted high availability requires more operational overhead than a managed service.

Helicone

Helicone
Helicone

Helicone is the fastest to integrate. Change one line - the base URL - and your existing OpenAI or Anthropic client logs everything automatically. No SDK imports, no instrumentation code.

Python
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
    base_url="https://oai.helicone.ai/v1",
    default_headers={
        "Helicone-Auth": f"Bearer {os.environ['HELICONE_API_KEY']}",
    },
)
# Works identically - requests are logged and forwarded

The dashboard shows cost per request, cost per user, cost per feature, and latency percentiles. Helicone also ships response caching and user-level rate limiting as one-liner header additions - both genuinely useful in production.

Pricing: Free (100k requests/month) - Pro $20/month - Enterprise custom.

What it does well: Zero instrumentation overhead. Cost dashboards are the best in the category. Fastest time-to-value of any tool here.

What it does not do: No prompt versioning or management. No multi-step agent trace trees. If you need to trace an agent across 6 model calls, Helicone sees 6 independent requests.

Braintrust

Braintrust
Braintrust

Braintrust is built around a different philosophy: evaluation is the primary workflow, and logging is how you feed it. The core primitive is the experiment - you define a dataset, run your LLM pipeline against it, score each result, and compare experiments over time.

Python
import braintrust
import anthropic

client = anthropic.Anthropic()

experiment = braintrust.init(
    project="code-explainer",
    experiment="claude-sonnet-v2-system-prompt",
)

for item in dataset:
    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=512,
        messages=[{"role": "user", "content": item["input"]}],
    )
    experiment.log(
        input=item["input"],
        output=response.content[0].text,
        expected=item["expected"],
        scores={"quality": score_fn(item, response.content[0].text)},
    )

experiment.close()

Once you have multiple experiments, Braintrust shows a comparison table: experiment A scored 0.72 average quality, experiment B scored 0.81. Drill into regressions to see exactly what changed.

Pricing: Starter free (5 users, 100k log rows/month) - Pro $200/month - Enterprise custom.

What it does well: The most mature evaluation workflow in the category. Prompt versioning, dataset management, A/B experiment comparison, and online scoring. The TypeScript SDK is as complete as the Python one.

What it does not do: Self-hosting is not supported. At $200/month, the Pro plan is significantly more expensive than alternatives.

Phoenix by Arize

Phoenix by Arize
Phoenix by Arize

Phoenix is the open-source observability project from Arize AI and the most OTEL-native tool here. If your team already uses OpenTelemetry for distributed tracing, Phoenix plugs directly into that pipeline. LLM spans look like any other OTEL span.

Python
import phoenix as px
from phoenix.otel import register
from opentelemetry.instrumentation.anthropic import AnthropicInstrumentor

# Start a local Phoenix server - no sign-up, runs in memory
session = px.launch_app()

tracer_provider = register(
    project_name="code-assistant",
    endpoint="http://localhost:6006/v1/traces",
)

# Auto-instrument the Anthropic client - zero manual span creation
AnthropicInstrumentor().instrument(tracer_provider=tracer_provider)

import anthropic
client = anthropic.Anthropic()

response = client.messages.create(
    model="claude-sonnet-4-6",
    max_tokens=512,
    messages=[{"role": "user", "content": "Write a binary search in Go"}],
)

Phoenix runs entirely locally during development - no API key, no network call, no data leaving your machine. It also ships pre-built evaluators for hallucination detection, RAG faithfulness, Q&A correctness, and toxicity.

Pricing: Open source (free, self-hosted) - Arize managed cloud: Developer $99/month, Team $499/month.

What it does well: OTEL-native design integrates into existing observability stacks. Local-first development mode is the fastest way to debug without touching any cloud service. Built-in eval templates save significant setup time.

What it does not do: No prompt management or versioning. The local mode loses data when the process stops - you need the cloud or your own storage backend to persist traces.

LangSmith

LangSmith
LangSmith

LangSmith is the observability platform from the LangChain team. If you are building with LangChain or LangGraph, it is the lowest-friction option - set two environment variables and every chain and agent run is traced automatically.

Python
from langsmith import traceable
import anthropic

client = anthropic.Anthropic()

@traceable(run_type="llm", name="code-review-agent")
def review_code(code: str, language: str) -> str:
    response = client.messages.create(
        model="claude-sonnet-4-6",
        max_tokens=1024,
        system="You are a senior code reviewer. Be specific and technical.",
        messages=[{"role": "user", "content": f"Review this {language} code:\n\n```{language}\n{code}\n```"}],
    )
    return response.content[0].text

result = review_code(my_code, "python")  # traced in LangSmith

For LangGraph agents, every node execution, tool call, and state transition appears as a separate span. The dataset and regression testing workflow is the closest alternative to Braintrust.

Pricing: Developer free (5k traces/month, 14-day retention) - Plus $39/month (50k traces, 90-day retention) - Enterprise custom.

What it does well: Zero-config for LangChain/LangGraph apps. The best multi-agent trace visualization - LangGraph's state machine diagrams render directly in the trace view.

What it does not do: Self-hosting requires the Enterprise plan. The free tier's 5k trace limit fills quickly on real traffic. Cost attribution dashboards are less detailed than Helicone.

Which Tool Should You Use?

  • Building with LangChain or LangGraph? LangSmith - zero config, and the agent trace visualization is unmatched.
  • Want LLM monitoring running in 5 minutes with no code changes? Helicone - change the base URL, done.
  • Need to self-host everything on your own infrastructure? Langfuse - mature Docker deployment, full feature set, MIT licensed.
  • Running evaluations and A/B testing prompts regularly? Braintrust - built exactly for that.
  • Debugging locally with no account or cloud dependency? Phoenix - runs in-process, no signup, data stays on your machine.
  • Using OpenTelemetry already? Phoenix - your LLM spans slot into the same pipeline as your other traces.
  • Regulated environment where vendor data access is a hard no? Langfuse self-hosted or Phoenix self-hosted.

Common Pitfalls

  • Logging without evaluating. A trace showing "200 OK, 341 tokens, 1.2s" tells you the model responded, not whether the response was accurate. Set up at least one eval scorer before going to production.
  • PII in traces. User content appears verbatim in traces. Implement input scrubbing before traces leave your application layer - especially if your trace store is accessible to anyone not cleared for user data.
  • Over-relying on LLM-as-judge. LLM scorers are fast but biased toward verbose responses. Mix in deterministic scorers (regex match, JSON schema validation, exact match) for any dimension where ground truth is available.

Working with LLM Observability Data

Several DevToolLab tools are useful for digging into observability payloads:

  • JSON Formatter - Clean up raw trace API responses before reading
  • JSON Diff - Compare two traces to spot what changed between a passing and failing run
  • JWT Decoder - Decode API keys or auth tokens without sending them anywhere external
  • Regex Tester - Build PII scrubbing patterns to sanitize prompts before they hit your trace store
  • cURL Command Generator - Build queries against the Langfuse or LangSmith REST APIs

Conclusion

The right tool depends on your constraints. For a team shipping fast, Helicone gets you cost visibility immediately. For eval-driven prompt development, Braintrust is purpose-built. For on-prem requirements, Langfuse self-hosted covers the full feature set. For LangChain teams, LangSmith is the obvious default.

All five tools share one thing: they surface failure modes that HTTP status codes will never show. Production LLM apps fail in ways that look like success to existing monitoring. The sooner you have trace-level visibility, the sooner you can stop guessing why quality degraded.

Related Posts

Best API Documentation Platforms in 2026

Mintlify, ReadMe, Stoplight, Redocly and Scalar priced on what a custom domain and white-labeling cost, plus Swagger UI, the open-source core under most.

By DevToolLab Team•

Best WAF and Bot Detection Tools in 2026

Cloudflare, DataDome, HUMAN Security and Akamai compared on real pricing and behavior, plus CrowdSec, the open-source WAF that costs nothing to self-host.

By DevToolLab Team•

Developer Tools Pricing Index 2026

161 published prices for 76 products in 10 categories, each checked on September 26, 2026, with a free CSV. The same workload costs up to 36x more.

By DevToolLab Team•