Lists / AI infra
Best LLM observability platforms for AI engineers (2026): What ChatGPT, Claude & Gemini Recommend
LangSmith is named in all 50 recorded AI answers and Langfuse is named first most often. 13 LLM observability platforms ranked, with who should pick each.
LangSmith is the platform AI models name most for this question. It appears in 50 of 50 recorded answers and is named first in 8 of 50. Langfuse is the one the models most often put first, in 25 of 50 answers. Start a shortlist with those two, then pick by how your team works.
The counts record which products ten AI models name for an AI engineer, not how well those products work.
LLM observability tools help teams monitor, trace and evaluate AI system behaviour in production.
TL;DR
- LangSmith is the one name no model leaves out, so it belongs on any shortlist.
- Langfuse is the default the models reach for most, and the pick for teams that must self-host.
- Arize is listed almost everywhere and led almost nowhere, which reads as a component pick more than a first choice.
- Confident AI’s count depends on which model you ask. The Perplexity models name it in every answer and cite its own guide first. Two OpenAI models never name it.
- Check who wrote the page behind any AI recommendation in this category. Every comparable guide captured for this query comes from a vendor or carries vendor sponsorship.
1. LangSmith
Pick LangSmith if you want traces, evals and human review in one workflow and want the product every model on the panel includes.
Measured: named 50 of 50 (LangSmith 100%), first 8 of 50 (LangSmith 16%), average position 2.56.
LangSmith is the only product on the panel that no answer leaves out, which is why it leads this list. Presence and priority split here. The models include it everywhere and lead with a different product in most answers. ChatGPT’s answer to the head question puts it up front: “Best default for most AI engineers: LangSmith.”
LangChain, which builds it, says LangSmith is not limited to the LangChain and LangGraph ecosystem.
Confident AI’s guide disagrees in part: it says observability depth drops outside the LangChain ecosystem.
Paolo Perrone’s guide on Medium says its Insights feature runs unsupervised topic clustering across your traces.
If your agents already run on LangGraph, the framework question doesn’t touch you. If they don’t, test trace depth on your own code.
Pros
- Holds a 5 of 5 count with each of the ten models
- Annotation queues let domain experts review, label and correct specific traces
- Works with any framework through its
traceablewrapper
Cons
- First in only 8 of 50 answers despite full presence
- Seat pricing at $39 a seat limits access for cross-functional teams, by Confident AI’s account
- Self-hosting is restricted to the Enterprise tier
- Does not replace Datadog or another APM for infrastructure triage, by LangChain’s own account
Pricing: Developer $0 a seat a month, Plus $39 a seat a month, Enterprise custom. Best for: production AI agent debugging, evals, review and monitoring.
2. Langfuse
Pick Langfuse if you need to self-host traces, prompts and evals under an open-source licence, because it is the default the models most often commit to.
Measured: named 48 of 50 (Langfuse 96%), first 25 of 50 (Langfuse 50%), average position 1.96.
Langfuse is the product the models commit to. It holds the best average position on the panel, and more answers open with it than with any other product. Claude, Claude Fable 5, Perplexity and Sonar Reasoning Pro each call it their default in at least one answer. ChatGPT is the one model below a full count, yet it calls Langfuse its default in some of the answers that do name it.
LangChain’s guide says the recent acquisition by ClickHouse creates a roadmap question buyers should consider.
Confident AI’s guide says Langfuse has no native alerting on quality degradation.
Perrone’s guide says its monitors fire warning and alert thresholds into Slack, webhooks or GitHub Actions.
The two guides can’t both be current on alerting, so check Langfuse’s own docs before you rely on either account.
Pros
- Every product capability is MIT-licensed, including tracing, evaluations, prompt management and the playground
- Traces, prompts, datasets, experiments and evals share one workspace
- Every model except ChatGPT names it in all five of its answers
Cons
- ChatGPT names it in 3 of 5 answers
- Self-hosting shifts storage, ingestion, upgrades and reliability work onto your team
- Scoring requires custom implementation, with no built-in evaluation metrics, in Confident AI’s comparison
- Leaves failure clustering to you, with grouping still on its roadmap
Pricing: Hobby $0, Core $29 a month, Pro $199 a month. The two captured guides disagree on the Enterprise tier. LangChain’s guide lists it at $2,499 a month. Confident AI’s guide lists Enterprise from $2,499 a year. Best for: self-hostable traces, prompts, datasets and evals.
3. Arize
Pick Arize if your ML engineers want Phoenix running locally beside their notebooks now and a hosted production path later.
Measured: named 47 of 50 (Arize 94%), first 1 of 50 (Arize 2%), average position 4.13.
Arize is counted under both Arize and Phoenix, its free, self-hostable tool. It sits on almost every shortlist and leads almost none. Its gap between being named and being named first is 92 points, the widest on the panel. That pattern reads as a component pick: the models name it beside a default more often than as the default. GPT-5.6 Luna’s answer to the engineer question is the rare exception: “My default recommendation: Arize Phoenix”.
LangChain’s guide describes Phoenix as the local, notebook-friendly tool and AX as Arize’s managed platform.
Perrone’s guide says scoring live traffic, monitors and alerting belong to paid AX, not free Phoenix.
Pros
- Phoenix runs as a single container with no Kubernetes, by Arize’s own account
- The OpenAI family names it in every answer (Arize 100%)
- Instrumentation written for local Phoenix carries over to AX as a configuration change
- Listed in 47 of 50 answers
Cons
- Led an answer once in 50
- Phoenix ships under the Elastic License 2.0, which bars offering it as a service
- Arize’s acquisition by Dynatrace leaves its future uncertain, in LangChain’s reading
Pricing: AX Free, AX Pro $50 a month, AX Enterprise custom. Best for: local AI observability with a production platform path.
4. Braintrust
Pick Braintrust if your team runs AI quality as datasets, experiments and scores, because it is the eval-first default the OpenAI models reach for.
Measured: named 46 of 50 (Braintrust 92%), first 7 of 50 (Braintrust 14%), average position 3.33.
Braintrust’s first mentions come from the OpenAI models. GPT-5.6 Sol’s answer to the engineer question opens: “For most AI engineering teams, I’d start with Braintrust.” GPT-5.6 Luna also names it the default in more than one answer. Its own site, braintrust.dev, is the second most cited host in the run, with 99 citations, and it is the first citation in most GPT-5.6 Sol answers.
LangChain’s guide says Braintrust starts from the eval loop. Its best fit, the same guide says, is a team that already thinks in datasets, scores and product-quality experiments.
Perrone’s guide says Braintrust is built for the eval loop and opinionated about it.
Pros
- Anthropic and Google families both name it in every answer (Braintrust 100%)
- Production failures convert into eval dataset rows in one click
- Runs its data plane inside your own cloud account on a commercial plan
Cons
- ChatGPT includes it in 3 of 5 answers
- No open-source build to read or fork
- Lacks deployment and an LLM gateway, by LangChain’s comparison
Pricing: Starter $0, Pro listed at $249 a month, Enterprise custom. Best for: eval-first teams.
5. Confident AI
Pick Confident AI if product managers and QA need to score and annotate production traces without waiting on an engineer.
Measured: named 34 of 50 (Confident AI 68%), first 6 of 50 (Confident AI 12%), average position 4.18.
Confident AI is counted under its own name and DeepEval, its evaluation framework. No product in the top five depends more on which model you ask. Claude, Claude Opus 5, Perplexity and Sonar Reasoning Pro name it in every answer. GPT-5.6 Sol and ChatGPT name it in none. Perplexity’s comparison answer describes it as “positioned as the evaluation-first leader”. Every model that cites Confident AI’s own guide names Confident AI in all five of its answers, which makes it the clearest case in this category of a vendor’s guide feeding the answers.
That guide ranks Confident AI first. It says product managers, QA and domain experts can run evaluation cycles without engineering involvement.
Pros
- Scores every trace, span and conversation thread automatically, by its own account
- Quality alerts fire through PagerDuty, Slack and Teams
- The Anthropic family names it in almost every answer (Confident AI 93.3%)
- Named first in 6 of 50 answers, ahead of every product below it
Cons
- Absent from every GPT-5.6 Sol and ChatGPT answer
- Cloud-based and not open source, though enterprise self-hosting is available
- May be more platform than lightweight tracing needs, its own guide concedes
Pricing: Free tier, Starter $200 a month with unlimited seats, Team $2,000 a month, custom Enterprise. Best for: cross-functional teams that need evaluation, alerting, drift detection and annotation open to the whole team.
6. Helicone
Pick Helicone if your main question is what your LLM API calls cost and how they behave, and you want that without deep SDK work.
Measured: named 28 of 50 (Helicone 56%), first 0 of 50 (Helicone 0%), average position 6.71.
Helicone is the most-named product that no answer puts first. The models name it as an addition, not a starting point.
LangChain’s guide describes it as a gateway-style tool that sits in the request path. That placement gives it request logging, latency tracking, cost tracking and caching with a relatively fast setup. It also means Helicone sees what passes through the gateway, not your full agent state.
Confident AI’s guide describes a one-line integration made by swapping the API base URL. The same guide says many teams run a gateway and an observability platform side by side.
Budget for Helicone as a second tool beside a tracing platform, not as a replacement for one.
Pros
- Claude Fable 5 names it in all five answers
- Open-source core under Apache-2.0
- Caching and automatic failover span multiple providers
Cons
- Zero first mentions across 50 answers
- Monitoring stops at the request, with no view into agent chains
- Now part of Mintlify, with maintenance-mode language LangChain says deserves roadmap diligence
Pricing: Free, Pro $79 a month, Team $799 a month. Best for: API-level visibility, cost tracking, caching and routing.
7. Datadog
Pick Datadog if your platform team already runs production in Datadog and wants LLM spans next to APM, logs and infrastructure.
Measured: named 21 of 50 (Datadog 42%), first 0 of 50 (Datadog 0%), average position 6.71.
Datadog enters the answers as a complement, never a default. Sonar Reasoning Pro, answering the engineer question, lists Datadog LLM Observability among the key complements for specific situations.
LangChain’s guide says its advantage is correlation, with AI spans beside APM traces and infrastructure metrics. That guide covers Datadog Agent Observability.
Confident AI’s guide covers Datadog LLM Observability. It prices that product per LLM request. It also says the product has no built-in evaluation metrics for output quality.
The two guides describe different Datadog products, so confirm which product and pricing unit applies before you compare quotes.
Pros
- GPT-5.6 Luna names it in 4 of 5 answers
- Adds no new vendor for teams already on Datadog
- Publishes span-based pricing for its Pro plan
Cons
- Gemini 3.5 Flash never names it
- Teams outside Datadog rarely adopt it for LLM observability alone, per LangChain
Pricing: free tier, with Pro from $160 a month billed annually for the first 100,000 LLM spans. Best for: AI telemetry inside an existing Datadog stack.
8. Comet
Pick Comet’s Opik if you want tracing, evals, prompt versioning and monitoring on your own hardware under an Apache 2.0 licence.
Measured: named 18 of 50 (Comet 36%), first 1 of 50 (Comet 2%), average position 4.89.
Comet is counted under Comet and Opik, its LLM observability platform. Its support sits in the Claude and Gemini models and is thin elsewhere. When a model does name it, it sits earlier in the answer than Helicone or Datadog.
Paolo Perrone writes in his Medium guide that he put Opik at the top of his list. Comet, which builds Opik, sponsored that guide. The guide says Opik’s Diagnostics feature groups the recurring silent errors it finds across traces.
Weigh that ranking knowing who paid for it, and weigh the feature description on its own terms.
Pros
- Apache 2.0 build ships the backend too, so the whole platform self-hosts
- Claude Opus 5, Claude and Gemini each name it in 4 of 5 answers
- A free cloud tier starts without a credit card
Cons
- GPT-5.6 Luna, ChatGPT, Perplexity and Sonar Reasoning Pro never name it
- Self-hosting Ollie, its trace-analysis agent, requires Enterprise
- Diagnostics scans consume Ollie tokens, so that cost scales with use
Pricing: free open-source version, free cloud tier, paid enterprise tier. Best for: teams who want tracing, evals, prompt versioning and monitoring on their own hardware.
9. MLflow
Pick MLflow if your ML team already runs its lifecycle on MLflow and wants LLM tracing in the same place.
Measured: named 15 of 50 (MLflow 30%), first 1 of 50 (MLflow 2%), average position 6.27.
No comparable guide captured for this query covers MLflow, so this entry rests on the panel alone. Claude Opus 5 and Sonar Reasoning Pro carry its count. Sonar Reasoning Pro frames it for teams already in the MLflow ecosystem, next to LangSmith for LangChain teams. Claude Opus 5 lists MLflow’s own article among the vendor guides that crown their publisher. mlflow.org is also one of the vendor-owned hosts the models cite, with 44 citations in the run. The count therefore reflects, at least in part, what MLflow publishes about itself. That is an inference from the citations, not a measurement of intent.
Pros
- Claude Opus 5 and Sonar Reasoning Pro each name it in 4 of 5 answers
- Led 1 of 50 answers, the same first count as Arize
Cons
- No OpenAI or Google model names it
- No pricing is recorded in the captured research
- No capability claims are recorded for it in the captured guides
Pricing: No pricing is recorded. Best for: Not recorded in the captured guides.
10. OpenObserve
Pick OpenObserve if you want LLM spans, logs, metrics and traces in one self-hosted OpenTelemetry backend.
Measured: named 14 of 50 (OpenObserve 28%), first 1 of 50 (OpenObserve 2%), average position 6.14.
OpenObserve’s count leans on one model. Claude Opus 5 names it in every answer, and notes in one of them that the OpenObserve blog ranks OpenObserve first. Sonar Reasoning Pro puts it forward when a team needs LLM and infrastructure observability together. openobserve.ai is a vendor-owned host with 27 citations in the run.
Confident AI’s guide calls OpenObserve an OpenTelemetry-native observability backend. It says OpenObserve stores logs, metrics, traces, frontend telemetry and LLM spans in one system.
That makes it an infrastructure team’s pick more than an evaluation team’s.
Pros
- Perplexity and Sonar Reasoning Pro each name it in 3 of 5 answers
- AGPL-3.0 self-hosting keeps deployment and data under your control
- Cloud pricing carries no per-host or per-seat charge
Cons
- Absent from every OpenAI and Google answer
- Evaluation, experiment, playground and MCP workflows are Enterprise features
- AGPL obligations may not fit teams modifying it for a proprietary service
Pricing: open-source self-hosting, with Cloud at $0.50 per GB ingested and $0.01 per GB queried. Best for: infrastructure-focused teams that want self-hosted LLM traces, logs, metrics and alerts in one backend.
11. Portkey
Pick Portkey if provider routing, fallbacks and retries are your hardest problem and request logs can come along with them.
Measured: named 14 of 50 (Portkey 28%), first 0 of 50 (Portkey 0%), average position 7.36.
Portkey is a gateway first, and the models treat it that way. No answer puts it first, and Gemini 3.5 Flash names it more than any other model.
LangChain’s guide describes it as primarily an AI gateway for routing, fallbacks, retries and guardrails. It says Portkey won’t replace a full eval workflow unless the team layers datasets, scoring and review on top.
Confident AI’s guide also calls it primarily a gateway, with observability as a built-in feature. It says teams that need to score outputs will need to pair Portkey with a dedicated platform.
Pros
- Gemini 3.5 Flash names it in 4 of 5 answers
- Cuts custom fallback, retry and routing code
- MIT-licensed, with a lightweight gateway footprint
Cons
- Never named by any OpenAI model
- Sits at an average position of 7.36, late in the answers that include it
Pricing: Developer free with 10,000 recorded logs, Production $49 a month, Enterprise custom. Best for: provider routing, fallbacks, guardrails and request logs.
12. Weights & Biases
Pick Weights & Biases Weave if your ML team already tracks experiments in W&B and wants LLM traces without another vendor.
Measured: named 9 of 50 (Weights & Biases 18%), first 0 of 50 (Weights & Biases 0%), average position 6.
Weights & Biases is counted under its own name, W&B and Weave. Its pattern mirrors Confident AI’s. The OpenAI models carry it, and Claude, Gemini, Gemini 3.5 Flash and Perplexity never name it. GPT-5.6 Luna’s comparison answer points to it for teams that already use a broader ML or experiment stack.
Confident AI’s guide says Weave, W&B’s tracing and evaluation product, adds LLM observability to the platform ML teams already use for training and experiments. The same guide calls Weave newer and less mature for production LLM observability.
Pros
- Each OpenAI model names it in 2 of 5 answers
- Carries W&B model versioning and artifact management into LLM work
- Structured trace capture comes with evaluation hooks
Cons
- Never named first
- No real-time quality alerting, per Confident AI’s guide
- Built for ML engineers rather than cross-functional teams
Pricing: Free tier, Teams $50 per seat per month, custom Enterprise. Best for: ML teams already on W&B that want LLM observability without leaving it.
13. Galileo
Pick Galileo if running LLM judges at production volume has become a cost line someone asks about.
Measured: named 8 of 50 (Galileo 16%), first 0 of 50 (Galileo 0%), average position 7.5.
Galileo is thinly but widely named. Models from all four families mention it, and none more than twice in five answers.
Perrone’s guide on Medium says Galileo’s differentiator is economic. It says Galileo distils LLM judges into compact Luna models to cut the cost of scoring traces. It adds that the cost figure is Galileo’s own benchmark, one to reproduce on your own traffic before you budget around it. Of the four automated-analysis tools it covers, the guide calls Galileo the only one that also goes after what the judging costs.
Pros
- Offline evals can be promoted to run as live guardrails
- Deploys as SaaS, in a VPC or on-premises
Cons
- Sits latest in the answers of any named product, at an average position of 7.5
- Listed as commercial rather than open source
- Zero answers put it first
Pricing: commercial, with SaaS, VPC and on-premises deployment. Best for: enterprises running evals at a volume where per-trace judge cost shows up as a line item.
How do the tools compare?
LangSmith, Langfuse, Arize and Braintrust are on almost every shortlist, and Confident AI is on most. The full record, with every answer, is at the LLM observability index.
| Rank | Vendor | Named | Answer share | Named first | First share | Avg position | Pricing model |
|---|---|---|---|---|---|---|---|
| 1 | LangSmith | 50/50 | LangSmith 100% | 8/50 | LangSmith 16% | 2.56 | Per seat, free Developer tier |
| 2 | Langfuse | 48/50 | Langfuse 96% | 25/50 | Langfuse 50% | 1.96 | Tiered, free Hobby tier |
| 3 | Arize | 47/50 | Arize 94% | 1/50 | Arize 2% | 4.13 | Free Phoenix, paid AX tiers |
| 4 | Braintrust | 46/50 | Braintrust 92% | 7/50 | Braintrust 14% | 3.33 | Free Starter, paid Pro |
| 5 | Confident AI | 34/50 | Confident AI 68% | 6/50 | Confident AI 12% | 4.18 | Free tier, flat monthly plans |
| 6 | Helicone | 28/50 | Helicone 56% | 0/50 | Helicone 0% | 6.71 | Free, then monthly tiers |
| 7 | Datadog | 21/50 | Datadog 42% | 0/50 | Datadog 0% | 6.71 | Span-based, free tier |
| 8 | Comet | 18/50 | Comet 36% | 1/50 | Comet 2% | 4.89 | Free open source and cloud, paid enterprise |
| 9 | MLflow | 15/50 | MLflow 30% | 1/50 | MLflow 2% | 6.27 | Not recorded |
| 10 | OpenObserve | 14/50 | OpenObserve 28% | 1/50 | OpenObserve 2% | 6.14 | Free self-hosting, usage-based cloud |
| 11 | Portkey | 14/50 | Portkey 28% | 0/50 | Portkey 0% | 7.36 | Free Developer, paid Production |
| 12 | Weights & Biases | 9/50 | Weights & Biases 18% | 0/50 | Weights & Biases 0% | 6 | Free tier, per seat |
| 13 | Galileo | 8/50 | Galileo 16% | 0/50 | Galileo 0% | 7.5 | Commercial |
The pricing-model column summarises the Pricing line in each entry, where the source for each price is marked.
LangSmith leads on presence. Langfuse leads on priority, with the most first mentions and the best average position. Below them, Arize and Braintrust complete a top tier that almost every answer includes. Confident AI stands alone in the middle. Everything from Helicone down is named in 28 of 50 answers or fewer, and only Comet, MLflow and OpenObserve are ever named first, once each.
The ranking follows presence. A buyer who wants to know which product a model would start with should read the named-first column instead, and that column starts with Langfuse.
Where do the models disagree?
The models agree on the leader and disagree below it. All ten have LangSmith as their most-named product, each at 5 of 5. The split shows how many of each model’s five answers named each vendor.
| Vendor | GPT-5.6 Sol | ChatGPT | GPT-5.6 Luna | Claude Opus 5 | Claude | Claude Fable 5 | Gemini | Gemini 3.5 Flash | Perplexity | Sonar Reasoning Pro |
|---|---|---|---|---|---|---|---|---|---|---|
| LangSmith | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| Langfuse | 5/5 | 3/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 |
| Arize | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 4/5 | 5/5 | 5/5 | 4/5 | 4/5 |
| Braintrust | 5/5 | 3/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 4/5 | 4/5 |
| Confident AI | 0/5 | 0/5 | 2/5 | 5/5 | 5/5 | 4/5 | 4/5 | 4/5 | 5/5 | 5/5 |
| Helicone | 2/5 | 0/5 | 4/5 | 3/5 | 3/5 | 5/5 | 3/5 | 4/5 | 2/5 | 2/5 |
| Datadog | 1/5 | 1/5 | 4/5 | 3/5 | 2/5 | 3/5 | 1/5 | 0/5 | 3/5 | 3/5 |
| Comet | 1/5 | 0/5 | 0/5 | 4/5 | 4/5 | 3/5 | 4/5 | 2/5 | 0/5 | 0/5 |
| MLflow | 0/5 | 0/5 | 0/5 | 4/5 | 3/5 | 2/5 | 0/5 | 0/5 | 2/5 | 4/5 |
| OpenObserve | 0/5 | 0/5 | 0/5 | 5/5 | 2/5 | 1/5 | 0/5 | 0/5 | 3/5 | 3/5 |
| Portkey | 0/5 | 0/5 | 0/5 | 1/5 | 3/5 | 2/5 | 3/5 | 4/5 | 1/5 | 0/5 |
| Weights & Biases | 2/5 | 2/5 | 2/5 | 1/5 | 0/5 | 1/5 | 0/5 | 0/5 | 0/5 | 1/5 |
| Galileo | 1/5 | 0/5 | 1/5 | 0/5 | 1/5 | 1/5 | 2/5 | 0/5 | 1/5 | 1/5 |
The sharpest split is Confident AI. GPT-5.6 Sol and ChatGPT never name it, while Claude, Claude Opus 5, Perplexity and Sonar Reasoning Pro name it every time.
Weights & Biases runs the other way. Each OpenAI model names it in 2 of 5 answers, and Claude, Gemini, Gemini 3.5 Flash and Perplexity never do.
Claude Opus 5 carries the long tail. It names OpenObserve in 5 of 5 answers and MLflow and Comet in 4 of 5, where the OpenAI models name OpenObserve and MLflow in none.
Gateways split by model too. Claude Fable 5 names Helicone in every answer and ChatGPT in none. Gemini 3.5 Flash names Portkey in 4 of 5 answers, and no OpenAI model names it.
The defaults differ as well. GPT-5.6 Sol names Braintrust as its default in most answers. GPT-5.6 Luna splits its defaults between Braintrust, LangSmith and Arize Phoenix. Claude and Claude Fable 5 name Langfuse as their default in most answers. Perplexity and Sonar Reasoning Pro pair Langfuse with Confident AI.
Why do the Perplexity models name Confident AI in every answer?
The likeliest reason is what they read. Every Perplexity and Sonar Reasoning Pro answer lists Confident AI’s own comparison guide as its first citation. confident-ai.com is the most cited host in the run, with 155 citations, and the top vendor-owned host.
That guide calls Confident AI the best LLM observability tool for evaluation and monitoring.
The same guide appears in Google’s organic results for the head question.
The pattern holds beyond Perplexity. Claude and Claude Opus 5 cite the same guide in some of their answers, and both name Confident AI in all five. GPT-5.6 Sol and ChatGPT never cite it and never name Confident AI. Their first citations mostly point to vendor documentation instead, such as docs.langchain.com and braintrust.dev. GPT-5.6 Sol makes Braintrust its pick in most answers.
Spotting a vendor guide does not remove the name. Claude Opus 5 flagged these guides directly: “Treat those rankings as ads.” It still names Confident AI in 5 of 5 answers.
The counts cannot prove cause. They are consistent with retrieval shaping the shortlist: a model names what its sources name, and in this category many sources are vendors writing about themselves. A product’s count measures visibility, and part of that visibility is the vendor’s own publishing. The totals in the table cannot show that. The citation trail can.
What should a buyer do with these counts?
Use the list as the set your colleagues will hear when they ask an AI the same question. LangSmith, Langfuse, Arize and Braintrust appear in nearly every answer, so expect to explain a choice outside that group.
Then pick by operating model, not by count.
LangChain’s guide says the right choice depends on the team’s operating model.
Perrone’s guide says Opik, Langfuse and Arize Phoenix all self-host for free.
On this shortlist, eval-first teams point to Braintrust or Confident AI. A team already on Datadog points to Datadog. Routing and cost problems point to Helicone or Portkey.
Open the citations behind any AI recommendation you rely on, and check whether the page was written by the vendor it recommends.
Finally, test on your own data. Perrone’s guide says an hour of your own traces will tell you more than a week of comparison tables.
How was the sample built?
10 models x 5 fixed prompts = 50 recorded answers, from the 2026-09 edition. Each model answered each prompt once. The five questions, verbatim:
- What is the best LLM observability and evaluation platform for an AI engineer? Name specific products.
- Which LLM observability and evaluation platform would you recommend to an AI engineer in 2026?
- Compare the top LLM observability and evaluation platform options right now.
- I’m an AI engineer and I need a LLM observability and evaluation platform. What should I use and why?
- Best LLM observability and evaluation platform for tracing and evals in production?
The panel covers four families. OpenAI gave 15 answers from GPT-5.6 Sol, ChatGPT (GPT-5.6 Terra) and GPT-5.6 Luna. Anthropic gave 15 from Claude Opus 5, Claude (Claude Sonnet 5) and Claude Fable 5. Google gave 10 from Gemini (Gemini 3.6 Flash) and Gemini 3.5 Flash. Perplexity gave 10 from Perplexity (Sonar Pro) and Sonar Reasoning Pro.
The run tracked 17 vendors and 13 were named. PromptLayer, Humanloop, Laminar and Sentry were tracked and never named, so they carry no rank.
Answer share is the share of recorded answers that named the product. Named first is the share of answers where it appeared before any other tracked product. Names are matched as strings, so Phoenix counts for Arize, Opik for Comet, DeepEval for Confident AI, and Weave for Weights & Biases. The full method is on the method page.
How does this list compare with the LLM observability guides?
The comparable guides rank products for purchase. The counts here record the names AI answers produce. Every comparable guide captured for this query comes from a vendor or carries vendor sponsorship.
LangChain’s guide compares LLM observability tools for production AI agents, with pricing, a best-for line and a pros-and-cons table for each tool. LangChain states its interest early: “We built LangSmith for teams that need observability to connect directly to the Agent Development Lifecycle”.
Confident AI’s guide is written by Jeffrey Ip, its co-founder. It puts Confident AI first in its list.
Paolo Perrone’s guide on Medium is the Comet-sponsored guide from the Comet entry above. It describes each platform from that platform’s own public documentation. Perrone writes that the section on where Opik will annoy you is his own.
The talk on the AI Engineer YouTube channel is given by Dat Ngo of Arize AI.
That talk walks through Arize’s own products and ranks nothing. The guides describe features, pricing and fit from documentation. None of them shows which products AI models actually name, how that changes by model, or which pages the models cite when they do. The panel adds exactly that: the per-model split, the gap between being named and being named first, and the citation trail behind a count.
What can these counts not tell you?
The counts measure presence in recorded answers. They do not measure product quality, reliability, support, adoption or fit with your stack. Being named is not the same as being recommended, because an answer can list a product only to warn against it.
Each model answered each prompt once, so a single answer can move a count. The data is one dated snapshot, the 2026-09 edition. Names are matched as strings, which can miss an unusual spelling. Answers came through model APIs, which can differ from consumer chat apps. The prompts were in English. Citation counts reflect the citations returned in the recorded API responses, and coverage varies by model.
Frequently asked questions
What are LLM observability tools?
They capture, visualise and evaluate what an LLM or agent does in production: every prompt, tool call, retrieval step and response, plus cost, latency and quality. They differ from ordinary application monitoring because the failures are usually semantic, such as a wrong answer or a bad tool choice with nothing thrown.
Confident AI’s guide splits the category into three camps: APM platforms adding LLM features, AI-native tracing tools and AI gateways.
What are the best observability tools for monitoring AI agents?
In the recorded answers, LangSmith, Langfuse, Arize and Braintrust are the tools the models name most, each in nearly every answer.
Agent monitoring needs the whole trace tree, one span per step, with inputs, outputs, latency and token spend for each.
LangChain’s guide is written for production AI agents. It recommends Langfuse to teams that want a self-hostable platform with traces, prompts, datasets and evals. It recommends Braintrust to teams that organise AI product quality around datasets, experiments, scores and eval runs. For framework-agnostic agent observability plus evals, monitoring and annotation queues, it recommends LangSmith, LangChain’s own product.
Is LangSmith only for LangChain users?
No. LangChain says LangSmith works with any framework or custom code.
Confident AI’s guide says the deepest integration is still with LangChain and LangGraph. Teams outside that ecosystem should test trace depth on their own code.
Which of these platforms are open source?
Langfuse’s core is MIT and Opik is Apache 2.0. Arize Phoenix ships under the Elastic License 2.0, which is source-available and a weaker promise than OSI open source. Braintrust, LangSmith and Galileo are commercial.
Confident AI’s guide lists Helicone as Apache-2.0, Portkey as MIT and OpenObserve as AGPL-3.0.
Should a team run one tool or several?
Many teams combine them. LangChain’s guide describes Datadog staying as the operations layer while a dedicated tool handles evals, and a gateway handling routing and spend beside a deeper platform. If you combine tools, decide where production failures are triaged and where eval datasets live.
Can a vendor pay to move up this list?
No. The publication does not sell rank or take affiliate money, and vendors cannot pay to appear, be reordered or be removed. The order comes from the recorded answers only.
Why are PromptLayer and Humanloop not on the list?
They were tracked and never named in any of the 50 answers. The same applies to Laminar and Sentry. A product that no model names gets no rank and no share.
How this list is ordered
The order is the measurement, not an assessment of the products. Answer share is the share of recorded answers that named the tool. Named first is the share where it appeared before any other tracked tool. Both are counts from one dated edition and are published in full on the category page.
A tool appears here only if it was named in the edition and its record carries a sourced claim. A product that was never named is not listed, and no position is sold.
Where to check it
- The full category record every answer, per-model split, cited sources
- The recorded answers raw output and counts
- The method how the panel runs and what is counted