Langfuse vs LangSmith vs Braintrust: which is named first in AI answers?
Langfuse was named first in 25 of 50 AI answers. LangSmith appeared in all 50 and led 8. Compare all three across five buyer questions and four model providers.
Figure 01 · Edition 2026-09
LLM observability and evals
How often each product was named, and how often it appeared first.
- LangSmithNamed: 100%Named first: 16%
- LangfuseNamed: 96%Named first: 50%
- ArizeNamed: 94%Named first: 2%
- BraintrustNamed: 92%Named first: 14%
- Confident AINamed: 68%Named first: 12%
- HeliconeNamed: 56%Named first: 0%
- DatadogNamed: 42%Named first: 0%
- CometNamed: 36%Named first: 2%
Top 8 by answer share · 50 answers
10 models · 5 prompts · Memetik Index, 2026-09
Shares use all 50 answers as the denominator. Multiple products can appear in one answer. First position does not measure recommendation strength.
Key findings
Langfuse is named first in 25 of 50 recorded answers, LangSmith in 8 and Braintrust in 7.
LangSmith appears in all 50 answers yet leads few of them, so mention order and first-position order are different measures.
Langfuse leads four of the five recorded questions, but LangSmith leads the direct comparison question.
Anthropic answers put Langfuse first 12 times in 15, while OpenAI answers favour LangSmith and Braintrust at 6 in 15 each.
One response per question-model pair makes this a descriptive snapshot, not a repeatability estimate.
You are comparing three LLM observability and evals tools, and the AI answers your buyers read name all three. What you need is the order. Langfuse is named first most often in this sample, leading 25 of 50 recorded answers against 8 for LangSmith and 7 for Braintrust, while LangSmith is the name most likely to appear at all.
Our LLM observability and evals research compares mention frequency and first position for Langfuse, LangSmith and Braintrust across recorded buyer questions. Langfuse was named in 48 of 50 answers and first in 25. LangSmith was named in 50 of 50 answers and first in 8. Braintrust was named in 46 of 50 answers and first in 7. The mention order and the first-position order are not the same.
Which of Langfuse, LangSmith and Braintrust is named first most often?
Langfuse. It takes first position in 25 of the 50 recorded answers, half the sample.
| Vendor | Named answers | Answer share | Named first answers | Named first share |
|---|---|---|---|---|
| Langfuse | 48/50 | 96% | 25/50 | 50% |
| LangSmith | 50/50 | 100% | 8/50 | 16% |
| Braintrust | 46/50 | 92% | 7/50 | 14% |
The September 2026 category contains 50 recorded answers from 10 models across 5 fixed English-language buyer questions. Answer share and first position are counted over that same set, so the two columns are directly comparable. Every figure links back to the raw answers in the Index.
Why does appearing in every answer not mean being named first?
Because the two counts ask different things. Answer share measures whether a tracked vendor appears. Named first records the earliest appearance among all 17 tracked vendors, including products outside the selected three.
A mention counts when the vendor’s name or a known alias appears in the answer text. LangSmith clears that bar in every one of the 50 answers. Clearing it says nothing about position in the paragraph.
First-position counts for the three selected products sum to 40 of 50 answers. Other tracked products occupy 10 first positions. So the first-position race runs across all 17 tracked names, including the 14 outside this comparison. Alias matching is string-based, and the method page sets out how the counts are built.
How does the leader change across the five recorded buyer questions?
The leader moves. Across the whole category, Langfuse led 4 of the 5 recorded questions, Braintrust led 1 of the 5 recorded questions, LangSmith led 1 of the 5 recorded questions. Ties are preserved as ties.
| Recorded question | Langfuse named first | LangSmith named first | Braintrust named first | All-category first-position leader |
|---|---|---|---|---|
| 1. What is the best LLM observability and evaluation platform for an AI engineer? Name specific products. | 3/10 (30%) | 1/10 (10%) | 3/10 (30%) | Braintrust, Langfuse |
| 2. Which LLM observability and evaluation platform would you recommend to an AI engineer in 2026? | 7/10 (70%) | 0/10 (0%) | 2/10 (20%) | Langfuse |
| 3. Compare the top LLM observability and evaluation platform options right now. | 3/10 (30%) | 5/10 (50%) | 0/10 (0%) | LangSmith |
| 4. I’m an AI engineer and I need a LLM observability and evaluation platform. What should I use and why? | 6/10 (60%) | 0/10 (0%) | 1/10 (10%) | Langfuse |
| 5. Best LLM observability and evaluation platform for tracing and evals in production? | 6/10 (60%) | 2/10 (20%) | 1/10 (10%) | Langfuse |
The comparison question is the outlier. Ask models to compare the top options and LangSmith takes first position 5 times in 10, while Braintrust takes it 0 times. Ask for a recommendation and Langfuse takes 7 of 10. Prompts stay fixed between editions, so these five splits can be tracked over time.
Do OpenAI, Anthropic, Google and Perplexity name a different tool first?
Yes. The provider families disagree sharply. In the 15 OpenAI answers, Langfuse was named first 2 times, LangSmith was named first 6 times, Braintrust was named first 6 times. In the 15 Anthropic answers, Langfuse was named first 12 times, LangSmith was named first 0 times, Braintrust was named first 0 times.
| Provider family | Recorded answers | Langfuse named first | LangSmith named first | Braintrust named first |
|---|---|---|---|---|
| OpenAI | 15 | 2/15 (13.3%) | 6/15 (40%) | 6/15 (40%) |
| Anthropic | 15 | 12/15 (80%) | 0/15 (0%) | 0/15 (0%) |
| 10 | 7/10 (70%) | 2/10 (20%) | 1/10 (10%) | |
| Perplexity | 10 | 4/10 (40%) | 0/10 (0%) | 0/10 (0%) |
Each sampled model is weighted equally. OpenAI contributes 15 answers; Anthropic contributes 15 answers; Google contributes 10 answers; Perplexity contributes 10 answers. If your buyers sit inside one assistant, the category total matters less than the row for that provider family.
What does this dated sample not tell you about these three tools?
It does not tell you how stable these counts are. One response was collected per question-model combination. The results are a descriptive snapshot. Repeated controlled trials would be needed to estimate repeatability or isolate wording effects.
The frame is also narrower than the category. The category tracks 17 products and 13 were named. This article selects Langfuse, LangSmith and Braintrust for comparison, so the other named products sit outside every table above. We report the dated sample and the exact questions so readers can judge whether a visibility comparison supports their conclusion. Every raw answer sits behind the Index entry for you to read.
Which evidence should you use to choose between Langfuse, LangSmith and Braintrust?
Use these counts for one decision only: how visible each name is in AI answers today. First position records where a product name appears in the text. Product quality, purchase suitability and recommendation strength are separate questions.
Your decision here is how to interpret an AI-visibility comparison. A product-purchase decision needs separate evidence about your requirements and the three products. We help software founders and growth teams understand their products’ measured visibility in AI answers, and you can follow the category through our research or subscribe for the next edition.
Frequently asked questions
How do LangSmith and Langfuse compare head to head in this sample?
LangSmith is named more often, Langfuse is named first more often. LangSmith was named in 50 of 50 answers, an answer share of 100%, and first in 8, a first share of 16%. Langfuse was named in 48 of 50 answers, an answer share of 96%, and first in 25, a first share of 50%. The gap between the two measures is the finding: total presence and lead position rank the pair in opposite directions.
How does Braintrust compare with Langfuse across the recorded answers?
Langfuse is ahead on both measures. Braintrust was named in 46 of 50 answers, an answer share of 92%, and first in 7, a first share of 14%. Langfuse was named in 48 of 50 and first in 25. The two draw level on one question. Asked for the best platform for an AI engineer with specific products named, Braintrust took first position 3 times in 10 and Langfuse 3 times in 10, and both are recorded as the leader for that question.
What are each tool’s visibility strengths and weaknesses in this sample?
LangSmith’s strength is coverage, Langfuse’s is lead position, Braintrust’s is provider-specific. LangSmith reaches every answer at 100% share but leads only 16%. Langfuse reaches 96% and leads 50%. Braintrust reaches 92% and leads 14%.
| Vendor | Named answers | Answer share | Named first answers | Named first share |
|---|---|---|---|---|
| Langfuse | 48/50 | 96% | 25/50 | 50% |
| LangSmith | 50/50 | 100% | 8/50 | 16% |
| Braintrust | 46/50 | 92% | 7/50 | 14% |
The weaknesses show up by provider family. Langfuse takes first position in only 2 of 15 OpenAI answers, while LangSmith and Braintrust take 0 of 15 Anthropic answers and 0 of 10 Perplexity answers each.
| Provider family | Recorded answers | Langfuse named first | LangSmith named first | Braintrust named first |
|---|---|---|---|---|
| OpenAI | 15 | 2/15 (13.3%) | 6/15 (40%) | 6/15 (40%) |
| Anthropic | 15 | 12/15 (80%) | 0/15 (0%) | 0/15 (0%) |
| 10 | 7/10 (70%) | 2/10 (20%) | 1/10 (10%) | |
| Perplexity | 10 | 4/10 (40%) | 0/10 (0%) | 0/10 (0%) |
Behind the figures
Method and sources
Edition 2026-09: 5 fixed buyer prompts per category, sent to 10 models through their APIs. Each category contains 50 recorded answers. Answer share is the fraction of answers naming a vendor. Named first records the vendor's position before other tracked vendors, not the strength of a recommendation. This is one dated sample; small differences can reflect answer variation. API responses can differ from consumer chat products. Read the full method.
- LLM observability and evals: recorded answers 50 answers · Edition 2026-09
Cite this research
Jeannie Wong. Langfuse vs LangSmith vs Braintrust: which is named first in AI answers?. Memetik, 11 September 2026. Data: edition 2026-09. https://www.memetik.ai/research/langfuse-vs-langsmith-vs-braintrust
Charts and figures: CC BY 4.0. Credit Memetik Index and link to the source.