The same question, put to ChatGPT, Claude, Gemini and Perplexity
Ask four AI assistants who to consider and you do not always get the same shortlist. These are the figures from just under 27,000 answers logged by LLM Scout, which I founded.
- Written by
- Frank Vitetta
- Published
- Last updated
- Subject
- Research
- Data from
- LLM Scout, which I founded
- Reading time
- 7 min read
In brief
- Over nine months, LLM Scout logged answers from ChatGPT, Claude, Gemini and Perplexity. It recorded whether each answer named the brand it was tracking for that prompt.
- On the 749 prompts that all four answered, they made the same call 72.8% of the time. Most of that was agreement to leave the brand out. None of the four named it on 53.7% of prompts. All four named it on 19.1%.
- On the other 27.2% they split. When the brand came up in at least one answer, all four named it only 41.2% of the time.
- For a business, one assistant's list is one opinion. Check a second, whether you are the supplier hoping to be named or the buyer building a shortlist.
Disclosure
Every figure in this note is from LLM Scout, which I founded. It monitors which brands AI assistants name in their answers, so I have a commercial interest in this subject. The data set is not public and nobody outside LLM Scout has reviewed it.
What is happening
When someone asks an AI assistant which suppliers to consider, the answer names some companies and leaves others out. I wanted to know how much that depends on which assistant they happen to ask. LLM Scout's tracking data gave me a way to look.
What was counted
For every prompt in the data, LLM Scout monitors one brand. I will call it the tracked brand. The mention rate is how often the tracked brand was named, as a share of responses. An answer full of other company names still counts as a miss if the tracked brand is not among them. Being named is also different from being recommended. An answer can name a company without endorsing it.
The data covers a limited set of tracked brands, 1,552 unique prompts and 26,996 responses. The prompts have buyer intent. They are the kind of question someone asks when choosing a product or a supplier.
The like-for-like result
The cleanest comparison uses only the prompts for which all four assistants returned a valid answer. There are 749 of them. For each one, I counted how many of the four named the tracked brand.
| Outcome for the tracked brand | Prompts | Share of 749 |
|---|---|---|
| None of the four named it | 402 | 53.7% |
| All four named it | 143 | 19.1% |
| Same call from all four | 545 | 72.8% |
| Exactly one named it | 117 | 15.6% |
| Exactly two named it | 45 | 6.0% |
| Exactly three named it | 42 | 5.6% |
| Split decision | 204 | 27.2% |
- None of the four Same call53.7%
- Exactly one Split15.6%
- Exactly two Split6.0%
- Exactly three Split5.6%
- All four Same call19.1%
Solid: same call from all four, 72.8% Hatched: split decision, 27.2%
All four assistants made the same call on 545 of the 749 prompts (72.8%). Most of that is agreement to leave the brand out. On 402 prompts none of the four named it. That means the tracked brand was absent. The answers may still have named other companies. On 143 prompts all four named it.
That leaves 204 prompts where the assistants split. The most common split is the starkest. On 117 prompts exactly one assistant named the brand and the other three did not.
The reading to avoid
It is tempting to summarise Table 1 as "the assistants agreed only 19% of the time". That is wrong. 19.1% is how often all four named the tracked brand. Agreeing that a brand does not belong in an answer is also agreement. Counting it gives 72.8%.
A fairer way to describe the disagreement is to set aside the 402 prompts where nobody named the brand. On the remaining 347, where at least one assistant did, all four named it on 143 (41.2%) and exactly one named it on 117 (33.7%). So when a brand appeared in any answer, more often than not (204 of the 347) at least one assistant left it out.
The headline rates, with a warning
Across the whole data set, the mention rates look like this.
Scroll sideways to see every column
| Assistant | Responses | Tracked brand named | Mention rate | 95% interval |
|---|---|---|---|---|
| ChatGPT | 8,555 | 3,078 | 36.0% | 35.0 to 37.0 |
| Claude | 6,583 | 1,816 | 27.6% | 26.5 to 28.7 |
| Gemini | 4,931 | 1,211 | 24.6% | 23.4 to 25.8 |
| Perplexity | 6,927 | 1,667 | 24.1% | 23.1 to 25.1 |
ChatGPT named the tracked brand in 36.0% of its responses, about one and a half times the rate for Gemini and for Perplexity. The gap between ChatGPT and Perplexity is 11.9 percentage points.
These rows are not like-for-like. The response counts run from 4,931 for Gemini to 8,555 for ChatGPT because each assistant answered a somewhat different mix of prompts. Some of the gap may come from that mix. The intervals are 95% Wilson intervals, which assume every response is independent. These are repeated runs of a limited prompt set, so the real uncertainty is wider than the table shows.
The like-for-like subset points the same way, though. Of the 117 prompts where exactly one assistant named the brand, that assistant was ChatGPT on 50, Claude on 26, Gemini on 21 and Perplexity on 20. ChatGPT accounts for 42.7% of those lone mentions, about twice any other assistant.
I do not know why ChatGPT is higher. Differences in training data are plausible. So are differences in how cautious an assistant is about naming companies. They are also speculation on my part. The data does not say.
Why it matters
The figures above are measurements. What follows is my reading of them.
If you are the supplier
A buyer who asks an assistant for options sees the companies it names and may never go looking for the ones it leaves out. In this data, which assistant answered changed whether the tracked brand appeared on 27.2% of the 749 prompts, a little over one in four.
That makes "we checked ChatGPT and we come up" a weaker statement than it sounds. It is a check of one assistant. In this data set that was the assistant most likely to name the tracked brand when the other three did not. A buyer using a different one may see a list without you on it.
If you are the buyer
The same finding applies to your own team. If one person researches suppliers with ChatGPT and another uses Claude or Gemini, they may be working from different shortlists without knowing it. Neither list is wrong. Each is a partial view. Nobody chose the differences between them.
The study checked one brand per prompt. Any two answers that differ on that brand are different lists. So 27.2% is a floor for how often the full lists differed on these prompts. It is not an estimate of that figure.
What to watch
Four working rules follow from this. None of them needs a tool.
- Check more than one assistant. To find out whether your company comes up, put the questions your customers would ask to all four. Do not stop at the one you use yourself.
- Record absence as well as presence. The most common outcome here was that no assistant named the tracked brand. Note where you are missing as carefully as where you are named.
- Repeat the check. All four assistants were updated during the nine months this data covers. A check is a snapshot of systems that keep changing.
- Treat an assistant's list as a starting point. For supplier research or a market scan, ask a second assistant the same question and compare the two lists. Then confirm the companies through sources you trust before anyone is ruled in or out.
Setting working rules like these for a team is part of the training I offer. It starts with a free 30-minute discovery call, described on the work with me page.
Method
A note built on LLM Scout's numbers should answer the questions I would put to anyone else's. Those questions are who ran it, on what, what was counted and what is missing.
- Who ran it
- LLM Scout, which I founded. This is not an independent evaluation.
- On what
- 1,552 unique prompts for a limited set of tracked brands, put to four assistants over nine months. The response counts in Table 2 add up to 26,996.
- What was counted
- Whether a response named the tracked brand for its prompt. The figures say nothing about where in the answer the brand appeared or whether the mention was favourable.
- How it was checked
- Every figure in this note was checked against the study tables on 5 October 2026. The intervals in Table 2 were computed from the counts.
- What is missing
- The data set is not public, so you cannot rerun the analysis. The brands and prompts are not named. This note does not cover how each assistant was queried or with what settings. Those choices can change an answer.
Sources
- LLM Scout: the origin of all the data in this note. I founded LLM Scout.
The figures were checked against the LLM Scout study tables on 5 October 2026. LLM Scout has not published the data behind them, which is why this list has no study link. No outside research is cited in this note.