Skip to main content
chudi.dev

I Asked 3 AI Engines the Same 16 Buying Questions. They Agreed Under a Third of the Time.

Published Chudi Nnorukam 11 min read

A 47-answer study of who ChatGPT, Claude, and Perplexity recommend in one B2B category. Six brands hold 89% of all mentions, the three engines overlap by about 30%, and a third of answers name nobody at all.

Why this matters

A 47-answer study of who ChatGPT, Claude, and Perplexity recommend in one B2B category. Six brands hold 89% of all mentions, the three engines overlap by about 30%, and a third of answers name nobody at all.

Three AI engines asked the same 16 buying questions agreed with each other under a third of the time. Across 47 answers they produced 134 brand mentions, six brands took 89% of them, and 16 answers named no vendor at all. If you are measuring your AI visibility on one engine, you are measuring roughly a third of the picture.

I ran this on my own category, AI visibility tooling, on 2026-08-13. I am a vendor in that category, so I am in the results, and I finished eighth. The method and the raw data are below so you can run it on yours and check mine.

The short answer (three findings)

One. The category is already consolidated. Six brands hold 89% of all 134 mentions. The long tail is not competing; it barely exists.

Two. The engines disagree sharply. On the questions where at least one of a pair named a vendor, Claude and ChatGPT overlapped 30%, Claude and Perplexity 28%, ChatGPT and Perplexity 40%. A single-engine score is not a category score.

Three. A third of the ground is unclaimed. 16 of 47 answers named no vendor at all. Those questions have no incumbent to displace.

Method

Sixteen questions, chosen as the things a buyer types in the week they are ready to spend. Not keywords. Pricing questions and “who does this” questions included on purpose, because those are where money is closest.

Each question was asked once on each of three engines, in a clean session with no personalization and web search enabled. Every brand named in the answer text was recorded, matched on brand name and common aliases as well as on the web address. Fifteen competitor brands plus my own were scored.

That gives 48 attempts. One Claude attempt failed with a tool error and is excluded rather than scored as an absence, because a failed request and a genuine non-mention are not the same event, and counting them together understates every brand in the table. That leaves 47 usable answers.

Share of voice is answers naming a brand, divided by usable answers. A brand named twice in one answer counts once.

BrandAnswers naming itShare of voice
Semrush2451%
Profound2349%
Otterly2349%
Peec AI1940%
Ahrefs1838%
Scrunch AI1226%
Athena49%
citability.dev (mine)36%
Similarweb24%
Conductor24%
Evertune24%
Writesonic24%

Those six leaders account for 119 of 134 total mentions. Everyone else is splitting scraps.

Two of the six are not AI visibility tools at all. Semrush and Ahrefs are established search tools that added AI features late, and they lead partly on the weight of a decade of corpus. That is worth sitting with if you are a specialist competing on product quality. The engines are not ranking product quality. They are reflecting how often and how credibly a name appears in the text they have read.

The engines do not agree

This is the finding I did not expect, and it is the one that changes how the measurement should be done.

PairAverage overlapQuestions compared
ChatGPT and Perplexity40%10
Claude and ChatGPT30%14
Claude and Perplexity28%14

Overlap is the shared brands divided by the combined brands, per question, averaged. It is only computed on questions where at least one engine in the pair named somebody, which is why the counts differ.

Roughly seven out of ten times, a brand one engine considers worth naming is not on the other engine’s list. That has a direct consequence: a vendor telling you your AI visibility score, measured on one engine, is telling you about one engine.

They also differ in temperament, consistently:

EngineUsable answersAvg brands named per answerDistinct brands ever namedAnswers naming nobody
Claude153.73111
ChatGPT162.62106
Perplexity162.2599

Claude names the most and refuses the least. Perplexity is the most conservative by a wide margin, declining to name any vendor in more than half its answers. Neither behavior is better. But if you measure on Claude alone you will think the category is crowded, and if you measure on Perplexity alone you will think it is empty.

A third of the answers name nobody

Sixteen of 47 answers, 34%, named no vendor at all. Nine of 16 questions had at least one engine come back empty.

One question came back empty on all three engines: “How do I find out why ChatGPT is not citing my website?”

Five more were empty on two of three, and three more on one of three. They are not obscure. How much does an AI visibility audit cost. How do I audit whether AI crawlers can access my site. Which consultants specialize in this. How can a company get recommended by ChatGPT. What is the best way to make content citable.

These are buying questions. They have intent, they have money behind them, and no brand currently answers them. The consolidation in the first table and the vacuum in this one are the same story told twice: the engines have a settled answer for “name some tools” and no answer at all for “help me with my specific problem.”

The full results

Every question, every engine, exactly as recorded.

#QuestionClaudeChatGPTPerplexity
1Best AI visibility tools for B2B SaaS in 2026?failedProfound, Peec, Otterly, Semrush, Ahrefs, ScrunchProfound, Peec, Otterly, Scrunch, Athena
2How do I track whether ChatGPT cites my brand?Profound, Otterly, Semrush, Ahrefs, SimilarwebSemrush, Conductornone
3Which tool monitors brand mentions in AI answers?Profound, Peec, OtterlySemrush, AhrefsProfound, Peec, Otterly, Semrush
4Best generative engine optimization software?Profound, Peec, Otterly, Semrush, AhrefsProfound, Peec, Otterly, Semrush, ScrunchProfound, Otterly, Semrush
5How can a SaaS company get recommended by ChatGPT?Semrushnonenone
6What is an AI visibility audit and who does them?Profound, Peec, Otterly, SemrushSemrush, Ahrefsnone
7Tools to measure AI citation share for a brand?citability.dev, Profound, Peec, Otterly, Semrush, AhrefsProfound, Otterly, Semrush, Ahrefs, ScrunchProfound, Peec, Otterly, Semrush, Ahrefs, Similarweb, Scrunch
8Who offers white-label AI citation reporting?Profound, Peec, Otterly, SemrushPeecnone
9How do I find out why ChatGPT is not citing my site?nonenonenone
10Best answer engine optimization tools?Profound, Peec, Otterly, Semrush, AhrefsProfound, Peec, Otterly, Semrush, AhrefsProfound, Otterly, Semrush, Ahrefs, Writesonic, Scrunch
11What tracks visibility across ChatGPT, Perplexity, Gemini?Profound, Peec, Otterly, Semrush, Ahrefs, Scrunch, EvertunePeec, Otterly, Semrush, Ahrefs, ScrunchProfound, Peec, Otterly, Semrush, Scrunch
12How much does an AI visibility audit cost?citability.dev, Profound, Otterly, Semrush, Ahrefs, Athenanonenone
13Alternatives to Profound for AI brand monitoring?Profound, Peec, Otterly, Semrush, Ahrefs, ScrunchProfound, Peec, Otterly, Semrush, Ahrefs, Writesonic, Scrunch, Evertune, AthenaProfound, Peec, Otterly, Ahrefs, Scrunch, Athena
14How do I audit whether AI crawlers can access my site?citability.devnonenone
15Which consultants specialize in AI search citation?Profound, Conductornonenone
16Best way to make content citable by LLMs?Ahrefsnonenone

What this study cannot tell you

I would rather state the limits than have you find them.

It is one category on one day. These numbers describe AI visibility tooling as three engines answered on 2026-08-13. Nothing here generalizes to your category without running it on your category.

One run cannot separate signal from noise. Each question was asked once per engine. Answer engines vary between runs on the same prompt, so a brand’s exact count carries run-to-run variation I have not measured. The large gaps are safe to read. A one or two mention difference between two brands is not. I did not compute confidence intervals, because with a single run there is nothing honest to compute them from.

Correction, added hours after publishing: I cannot yet separate “the engines disagree” from “the engines are unstable.” I went looking for prior work on this exact question and found the check I should have run first. Researchers at the University of St. Gallen asked the same prompt of the same engine up to ten times inside a 24 hour window and measured how much the named brand set changed. It changed a lot: Jaccard overlap of 0.33 to 0.48 depending on the category (Schulte, Bleeker and Kaufmann, “Don’t Measure Once: Measuring Visibility in AI Search (GEO)”, April 2026, four engines, 45 day window). My cross-engine overlap of 0.28 to 0.40 sits inside that band.

So one engine asked twice may disagree with itself about as much as two engines disagree with each other, and my design cannot tell the difference, because I asked each question once. What survives is the practical conclusion, and it survives either way: a single-engine, single-run visibility score is not a category score. What does not survive is my attributing that specifically to differences between the engines. Treat the overlap table as a measurement of instability, not of disagreement, until the next run separates them.

I am leaving the original numbers and the original wording above exactly as published, rather than quietly editing them to match what I now know. The second run is designed to settle this: each question asked several times per engine, so the within-engine baseline can be computed first and the cross-engine gap measured against it.

The question set is mine. I chose 16 questions I believe buyers ask. A different 16 would move the numbers. The full list is in the table above so you can judge that yourself.

Brand matching is text matching. A brand named in a way I did not anticipate would be scored absent. I matched names and aliases rather than web addresses alone, because models say “Profound” far more often than they print a web address, and address-only matching would have made almost everyone look invisible.

Being named is not being recommended. Question 13 asks for alternatives to Profound, and Profound is named in the answer. Presence in the text is what I counted.

What I take from it

I am in this category and I placed eighth with three mentions out of 47, all three from one engine, none of them on the general “best tools” question. That is the honest state of my own visibility, and publishing it is more useful to me than an estimate would be.

The strategic read is that the crowded questions are settled and the empty ones are not. Displacing a brand that appears in half of all answers is a multi-year project. Being the first real answer to a question no brand answers is a page. I know which of those I am doing next, and I have written the first one on why ChatGPT is not citing your website, the question that came back empty on all three engines.

I will rerun these 16 questions on the same three engines and publish the comparison, including if my number does not move.

If you want this run on your category rather than mine, that measurement is what I built citability.dev to do.

Data: 16 questions, 3 engines, 48 attempts, 47 usable answers, 134 brand mentions, collected 2026-08-13. Every result is in the table above.

· Sources & further reading

Sources & Further Reading

Further reading

What do you think?

I post about this stuff on LinkedIn every day and the conversations there are great. If this post sparked a thought, I'd love to hear it.

Discuss on LinkedIn