Quick Answer: Choose AI agent observability tools by trace depth; the field splits into agent-tracing platforms, APM extensions and LLM gateways. LangSmith, Langfuse, Arize Phoenix, Braintrust and Datadog Agent Observability record every model call, tool call and retrieval with tokens, cost and latency. Gateways like Helicone and Portkey log model requests; tool steps appear only under a shared session or trace ID.
One agent run can call a model six times, hit three tools and query a vector store. When the answer is wrong, a request log shows six successful API calls and nothing else. AI agent observability records the whole path, so you can see which step went wrong.
This guide groups the tools by type, compares what each vendor documents today, and shows how to set up AI agent monitoring before real traffic arrives.
What is AI agent observability?
AI agent observability is the practice of recording each agent run as a trace: one span per model call, tool call, retrieval and decision, with inputs, outputs, tokens, cost and latency attached.
A good setup answers three questions: what did the agent do, why did it fail, and is quality slipping? The first needs the full trajectory, including the arguments passed to each tool. The second needs errors, retries and retrieved documents on one timeline. The third needs scores and user feedback attached to traces over time.
How it differs from APM
Application performance monitoring (APM) tells you a request was slow or failed. An agent can return HTTP 200 in two seconds and still give a wrong answer or loop five times. The failure is semantic, so you need content and sequence, not just timings. The two still coexist: an OpenTelemetry Collector can send one stream of spans to your APM backend and an LLM-specific tool at once.
Gateway logs are not agent traces
An LLM gateway sees every model request. It does not see the tool calls and retrieval steps between them unless your code passes a shared trace or session ID. Gateway logging helps with cost and access control; it doesn't replace tracing inside the agent.
Which AI agent observability tools do production teams use?
Production teams mostly choose from three types: agent-tracing and evaluation platforms (LangSmith, Langfuse, Arize Phoenix, Braintrust), APM platforms that added agent tracing (Datadog Agent Observability), and LLM gateways that log every request (Helicone, Portkey).
| Tool | Type | Open source or self-host | Agent tracing depth | Evals built in | OpenTelemetry ingest | Cost tracking |
|---|---|---|---|---|---|---|
| LangSmith | Tracing and eval platform | Enterprise tier (self-hosted add-on) | Yes (model, tool and retrieval runs) | Yes (offline and online) | Yes (OTLP endpoint) | Yes (from token counts) |
| Langfuse | Tracing and eval platform | Yes (open source, Docker; some add-ons licensed) | Yes (agent, tool, retriever types) | Yes (online and offline) | Yes (OTLP endpoint) | Yes |
| Arize Phoenix | Tracing and eval platform | Yes (Elastic License 2.0, free to self-host) | Yes (model calls, retrieval, tool use) | Yes (code and LLM-as-a-judge) | Yes (built on OpenTelemetry) | Yes |
| Braintrust | Eval-first platform | Enterprise tier (self-hosted data storage) | Yes (llm, function and tool spans) | Yes (offline and online) | Yes (OTLP exporter) | Yes (per LLM span) |
| Datadog Agent Observability | APM extension | Not documented | Yes (agent and workflow steps, tool calls) | Yes (quality, privacy, safety) | Yes (GenAI conventions 1.37+) | Yes (estimated per request) |
| Helicone | LLM gateway | Yes (Apache 2.0, Docker or Kubernetes) | Session-level (LLM, vector and tool calls by session ID) | No (stores scores from other frameworks) | Via OpenLLMetry async logging | Yes |
| Portkey | LLM gateway | Enterprise tier (data plane in your VPC); gateway open source | Trace-level (trace IDs; LangGraph and CrewAI auto-instrumentation, beta) | Cookbook for custom batch evals | Yes | Yes |
Capabilities as documented by each vendor on 25 September 2026; links in the text.
Tracing and evaluation platforms
LangSmith's documentation says every unit of work an agent performs, such as a model call, tool invocation or retrieval, is recorded as a run, and runs for one operation form a trace. Its evaluation docs describe adding production traces to a dataset so a failure you saw once becomes a test you run every time. Langfuse, Phoenix and Braintrust document the same loop: score live traces, save interesting ones as datasets, and rerun them before each release.
APM with agent tracing
Datadog Agent Observability puts agent traces next to your service traces and monitors, with a span for each agent choice or workflow step and evaluations for quality, privacy and safety.
Gateways
Helicone's sessions group related LLM calls, vector queries and tool calls under one ID; its docs say it doesn't run evaluations but stores scores from other frameworks. Portkey groups LLM calls under a trace ID you pass in, auto-instruments LangGraph and CrewAI in beta, and accepts OpenTelemetry data from the rest of your stack.
Which one fits
- Choose LangSmith for the trace-to-dataset-to-eval loop in one hosted product, if you already build on LangChain or LangGraph.
- Choose Langfuse or Arize Phoenix when trace data must stay on your infrastructure without an enterprise contract.
- Choose Datadog Agent Observability when your services already report to Datadog.
- Choose Braintrust for evaluation-heavy teams.
- Choose a gateway when cost control matters more than step-level debugging.
If evals matter most, compare AI agent evaluation tools separately.
How do open-source and commercial observability tools compare?
Open-source and self-hostable tools keep prompts, outputs and tool results on infrastructure you control, at the price of running it yourself. Hosted commercial tools remove the operations work but store your trace data with the vendor.
Traces hold the same sensitive data as the application, so data residency is the first question.
| Deployment model | Where trace data lives | Who runs it | Documented examples |
|---|---|---|---|
| Open-source self-host | Your servers or cloud account | Your team | Langfuse, Arize Phoenix, Helicone |
| Enterprise self-host or hybrid | Your cloud account; vendor may run the control plane | Shared | LangSmith, Braintrust, Portkey |
| Vendor-hosted | Vendor cloud | Vendor | Datadog Agent Observability; hosted tiers of the others |
Langfuse is open source and can be self-hosted with Docker, though some add-ons need a license key. Arize releases Phoenix under the Elastic License 2.0, free to self-host with no feature limits, and sells Arize AX as a managed enterprise platform.
Self-hosting costs you a database, storage and upgrade time. Hosted tools meter usage instead; Datadog, for example, meters Agent Observability on LLM spans ingested.
What should an observability tool log for every agent run?
Log enough to replay the run's decisions: identifiers, versions, every step's inputs and outputs (redacted), tool arguments and results, retrieved documents, tokens, cost, latency, the outcome and any feedback.
Teams moving an agent from pilot to production, with or without an implementation partner, should make tracing a launch requirement. Many failure modes in AI implementation appear only under real traffic, and without traces you can't tell a bad prompt from a bad tool.
| Field | Why it matters |
|---|---|
| Trace ID, run ID, session and pseudonymous user ID | Stitch multi-turn conversations and find one user's runs |
| Agent version, prompt version, model and parameters, tool schema version | Explain why behavior changed after a release |
| Inputs and outputs per step, redacted | See what the model saw and said without storing raw PII |
| Tool name, arguments, result, error and retries | Wrong arguments and unhandled tool errors are common agent failures |
| Retrieved document IDs and scores | Separate retrieval misses from reasoning mistakes |
| Tokens in and out, cost, latency per span and per run | Catch loops and cost spikes, and find slow steps |
| Final outcome, user feedback and eval scores | Detect quality decline across runs, not just single failures |
To catch quality decline:
- Run online evaluators on sampled production traces.
- Alert when scores or thumbs-down rates move.
- Route flagged runs to a reviewer.
How does OpenTelemetry support agent observability?
OpenTelemetry's GenAI semantic conventions give every tool the same names for agent spans, attributes and metrics, so you instrument once and send the traces to any backend that reads them.
The conventions now live in their own repository, and the OpenTelemetry GenAI semantic conventions cover events, exceptions, metrics, model spans and agent spans. Operation names include create_agent, invoke_agent, invoke_workflow, execute_tool, retrieval and plan. Attributes such as gen_ai.agent.version, gen_ai.request.model and gen_ai.usage.input_tokens carry the context.
Their status is still Development, so names can change between releases. Pin the convention version your instrumentation emits. Datadog ingests traces that follow the OpenTelemetry 1.37+ GenAI conventions or OpenInference, and Arize Phoenix is built on OpenTelemetry and OpenInference. Langfuse accepts OTLP traces and maps attributes because the conventions are still evolving, and LangSmith, Braintrust and Portkey also document OpenTelemetry ingest. Helicone logs traces sent through OpenLLMetry.
The payoff is portability: change backends, or run two, without re-instrumenting.
How do you track agent cost and latency over time?
Attach tokens, cost and latency to every model and tool span, roll them up per run, then chart them per task type, per user or tenant and per model, with alerts on cost per completed task and p95 run latency.
Cost per call misleads: a cheap model that needs four retries can cost more per resolved task than a pricier one. Track cost per successful outcome and step count per run, where loops show up first.
Most tools compute cost from token counts and a model price table. Datadog estimates cost per LLM request from public pricing and token counts, Helicone calculates it exactly through its gateway, and LangSmith lets you edit model prices, which matters for self-hosted models.
For latency, split time to first token (the OpenTelemetry metric gen_ai.client.operation.time_to_first_chunk), total model time and tool time. Alert on:
- Runs above a step or token ceiling
- Cost per task above a rolling baseline
- Tool error or retry rates above normal
- p95 latency regressions after a deploy
What mistakes should you avoid when monitoring agents in production?
The costly mistakes in AI agent monitoring are logging raw PII, sampling away failures, skipping version tags and building dashboards nobody is paged from.
- Logging PII unredacted. Prompts and tool results carry customer data. Redact at the SDK or Collector before export, and keep message capture opt-in.
- Sampling away failures. Head sampling at 10% drops 90% of your errors too. Keep every errored or low-scored run, and sample only successes.
- No version tagging. Without agent, prompt and model versions on each trace, you can't tell which release broke behavior.
- Tracing only the LLM calls. Request logs alone miss tool and retrieval steps unless they share a session or trace ID.
- Dashboards without alerts. Tie every alert to an owner.
- Ignoring trace limits. LangSmith caps a trace at 25,000 runs; long-running agents need threads or sessions.
How Origins AI monitors agents in production
Origins AI (originshq.com) builds and deploys AI agents for product teams through its AI workflow and agent development services, whose page lists encryption at rest and in transit, secure authentication, continuous security monitoring and least-privilege access as its data-security controls.
For model traffic from engineering tools, the Origins AI Coding Tool pairs a self-hosted LLM gateway with an AI code audit server. According to its product page, the gateway logs every request, token count and response in your environment, tracks usage by team, project and engineer, and redacts secrets and PII before content reaches the model. In on-premise and air-gapped deployment modes, no source code is sent to any external service; in hybrid mode, the code context submitted to a hosted model leaves your network.
For process agents, the Origins AI Agentic Automation page describes setting autonomy boundaries, escalation triggers and approval thresholds before a one-process pilot, then tracking performance metrics as it scales. For voice agents, the Origins AI Voice AI page states that every conversation, escalation and action is logged with timestamps, user IDs and session context.
Startups hiring a partner for production-ready agents should ask to see the trace and cost view before launch.
Talk to an engineer
Planning an agent launch and want tracing, cost tracking and alerts in from day one? Book a call with an engineer.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


