Perplexity Named Zero Brands, 40 Answers in a Row. Your Category Is Not the Reason.
480 answers across ChatGPT, Claude and Perplexity, in a two-year-old category and a twenty-year-old one. Which engine goes quiet is a property of the engine.
Why this matters
If you check your AI visibility once and see a gap, you are probably looking at noise. An engine asked the same question twice agrees with itself only about two thirds of the time. Any decision you make from a single check is a coin flip wearing a lab coat.
Earlier today I asked three AI engines the same 16 buying questions, once each, and published what came back. Perplexity named nobody on every question that was not a straight “list the tools” request. I could not tell you why. Two explanations fit equally well: the category is young and thin, or that is simply how Perplexity answers.
So I ran it again, properly. Same 16 question shapes, asked five separate times per engine instead of once, in two categories at the same time: AI visibility tooling, which is about two years old, and project management software, which has been settled since the 1990s. 480 answers. Every call completed; none were dropped.
Here is what came back.
The short answer
Perplexity going quiet is a property of Perplexity. It named nobody on 100% of non-list questions in the young category and 75% in the twenty-year-old one. Claude sat at 30% and 35%. The gap between the engines barely moves when you swap the entire market underneath them.
The one large public dataset says the opposite. It reports Perplexity as the engine least likely to decline. I measured it as the most likely, in both categories, by a wide margin. I can name two reasons we may be counting different things, and I still cannot close the gap.
Question shape does almost all the work. Ask “best X tools” and every engine names four or five brands. Ask literally anything else and between a third and all of the answers name nobody. That split is far larger than the difference between a new market and an old one.
The two-year-old category is more concentrated than the twenty-year-old one. Six brands hold 90% of all mentions in AI visibility. In project management, six brands hold 83%. I expected the opposite and was wrong.
And I was wrong about something I published this morning. More on that below, because it is the finding that changed my own plans.
If you want the actions rather than the evidence, skip to three things to do this week. Everything between here and there is how I know they are worth doing.
What I changed, and why it matters
Run 1 asked each question once per engine. That design cannot tell “the engines disagree with each other” apart from “each engine is unstable and gives a different answer every time.” I said so in a correction a few hours after publishing, once I found the researchers who had already caught it.
Run 2 fixes it directly. Asking each question five times per engine gives a baseline: how much does one engine disagree with itself across repeats? Only once you have that number can you say whether the disagreement between engines is bigger than the noise. Same measurement on both sides, so the two numbers are comparable.
Adding a control category does the same job for the other explanation. If AI visibility tooling is thin ground where models have little to say, then running the same 16 question shapes against a market with two decades of reviews, comparisons and buyer guides should light everything up. Any effect that survives both categories is about the engines.
This is not a free thing to run, and it is worth saying what it costs so you can judge whether to do it yourself. 480 separate calls, two hours and twenty minutes of unattended machine time, two categories held side by side so neither one gets a better day than the other, plus a rewrite and a restart partway through when I found a flaw. Asking once is cheap. Asking enough times to believe the answer is the whole expense.
I also removed a subtle bias partway through the run, and it is worth naming because it would have flattered me. The engines had a three minute cutoff. Claude’s answers were landing at two to two and a half minutes, with the slowest ones brushing that ceiling. The slow answers are the ones where the model searched hard, and those are exactly the answers most likely to name somebody. A cutoff there does not drop a random answer, it drops a full one, which pushes the “named nobody” rate up in precisely the direction my headline claims. I raised the limit to seven minutes and restarted. Fixing that was cheaper than disclosing it.
The engines do not behave the same way, and swapping markets barely changes it
Non-list questions only. Each cell is 40 answers.
| Engine | AI visibility (young) | Project management (settled) |
|---|---|---|
| Claude | 30% named nobody | 35% named nobody |
| ChatGPT | 75% named nobody | 50% named nobody |
| Perplexity | 100% named nobody | 75% named nobody |
Read down the columns rather than across. In both markets the ordering is identical and the spread is enormous: Claude names somebody on roughly two thirds of these questions, Perplexity on a quarter at best.
Now read across. Moving from a two-year-old market to a twenty-year-old one buys ChatGPT 25 points and Perplexity 25 points. It moves Claude by five points in the wrong direction. So category maturity is real for the two search-grounded engines and close to absent for the one that leans on what it already knows. But it is a smaller effect than the difference between the engines themselves, and it never reorders them.
One judgment call is buried in that table: which questions count as “list” questions. One of the sixteen, “who offers white-label reporting for agencies,” genuinely reads both ways. Counting it the other way gives Claude 34% and 34%, ChatGPT 80% and 43%, Perplexity 100% and 71%. Same story, so the classification is not carrying the result.
What the published numbers say
I want to be careful here, because there is a real contradiction and I do not think I get to resolve it by myself.
MaxAEO published an analysis of 11,520 answers across these same three engines, collected between April and July 2026. Their deferral rates run Perplexity 1.3%, ChatGPT 2.1%, Claude 9.4%. That is my ordering exactly inverted, and it is not close: they have Perplexity declining once in 77 answers, and I measured it declining three times in four.
Both cannot describe the same thing. After reading their write-up properly I can name two reasons for the gap, though neither settles it.
The first is that we are not counting the same event. They define a deferral as an answer that “names zero vendor brands and returns only evaluation criteria or clarifying questions.” I counted every answer that named zero brands, whatever else the answer did. So an answer that discusses the topic at length and never names a vendor is empty by my count and is not a deferral by theirs. That is the exact shape most of Perplexity’s answers took in my run, which means this difference alone could account for a large part of the distance between us.
The second is the sample. Their 320 prompts cover eight categories: analytics, CRM, security and compliance, HR and payroll, developer infrastructure, customer support, data pipelines, and marketing automation. Every one of those is a settled market. There is no young category in their study, so their sample sits much closer to my control condition than to my test condition. Their own numbers also move a great deal by category, with Claude’s deferral running 21.5% in security and compliance against 9.4% overall.
What I can say is that on 240 answers in a settled market, using a question set I have published in full, the ordering came out backwards from theirs, and I would rather write that down than quietly pick the study that agrees with me.
This is not a thin research area, and I did not find it early enough last time. Worth reading alongside this:
- Semrush’s study with Kevin Indig (the ghost citations study) ran 115 prompts and found comparative prompts produce a 43.3% brand mention rate against 18% for informational prompts. That is the same shape effect I am measuring, from a different angle, on four engines that do not include Claude or Perplexity.
- Peec AI’s analysis of roughly 200,000 responses across eight engines attributes category-level variation to weaker model priors in fragmented markets. My control category is the test of that idea, and it moved two engines out of three.
- A team publishing as arXiv 2606.23057 ran 3,750 API calls with five repeats per query across GPT-5.2, Gemini 3 Flash and Perplexity, and found all three agreed fully on only 41.6% of queries. They put their dataset on Zenodo, which is the reason I can tell you that.
- Schulte, Bleeker and Kaufmann at the University of St. Gallen (April 2026) measured how much one engine disagrees with itself inside 24 hours, and got brand-list overlap of 0.33 to 0.48. That is the paper that caught my run 1 error.
The engines really do disagree, and now I can prove it
This is the correction from run 1, closed out.
Take any two answers and measure how much their brand lists overlap, on a scale where 0 means no shared brands and 1 means identical lists. Score only the pairs where at least one side actually named somebody, since two empty answers agreeing about nothing is not agreement.
| Same engine, asked again | Two different engines | |
|---|---|---|
| AI visibility | 0.629 | 0.397 |
| Project management | 0.677 | 0.425 |
An engine asked the same question twice agrees with itself about 63% to 68% of the way. Two different engines agree about 40% to 43% of the way. The gap between those is the real cross-engine disagreement, and it holds in both markets.
So run 1’s conclusion survives, but only now does it have the right support under it. Both facts are true at once and both matter: these engines are genuinely unstable run to run, which the St. Gallen paper found first, and they genuinely disagree with each other beyond that instability. A visibility score from a single engine on a single day is measuring two different kinds of variation at once and reporting them as one number.
What that means if you are measuring your own visibility
That 63% is not a statistic about engines. It is a warning about your own reporting, so let me put it in the terms that actually cost money.
One measurement is not a fact. If you ask an engine where you stand and it names you, ask again tomorrow and there is roughly a one in three chance the answer changes. Not because anything about your site changed. Because that is how these systems behave.
A gap you saw once is probably not a gap. This is the mistake I made in run 1 and it is the expensive one. I found eight questions where nobody was named, treated them as open ground, and built a plan on it. Fifteen tries per question later, seven of those eight had owners the whole time. If you are choosing what to write next based on a single check, you are picking targets out of noise about half the time.
Movement is only visible across repeated measurements. You cannot tell whether a change you made worked by measuring once before and once after. Both numbers carry the same one-in-three wobble, so a real improvement and a lucky reading look identical. The only way to see a trend is to measure the same questions repeatedly over time and watch where the middle of the range goes.
None of that is an argument for a fancier tool. It is an argument for asking more than once, whoever does the asking. If you are running this by hand, ask five times and only believe results that show up in at least three of them. That single rule would have saved me from the mistake in run 1.
Question shape beats category age
The split that dominates everything else, in the young category:
| Engine | “Best X tools” questions | Everything else |
|---|---|---|
| Claude | 8% empty, 4.97 brands per answer | 30% empty, 2.05 brands |
| ChatGPT | 13% empty, 4.90 brands | 75% empty, 0.40 brands |
| Perplexity | 13% empty, 4.47 brands | 100% empty, 0.00 brands |
And in the settled category, list questions run 13% to 15% empty at 3.40 to 3.75 brands per answer, while everything else runs 35% to 75% empty at 0.75 to 1.43 brands.
Notice that the list-shaped questions look nearly identical across both markets. A twenty-year-old category does not produce longer brand lists than a two-year-old one when you ask directly. What the old category buys you is better answers to the indirect questions, and only on the two engines that search.
The practical read: if your visibility strategy is aimed at “best tools” listicles, you are competing on the one question shape where every engine already has a confident answer and six incumbents already own it. The questions where an engine currently names nobody are the ones where there is room, and most of them are not list-shaped.
The young category is more concentrated, not less
I expected a new market to be scattered and an old one to be consolidated. It is the other way round.
AI visibility, 240 answers, 672 mentions, 13 distinct brands named: Otterly 17.9%, Semrush 17.1%, Profound 17.0%, Peec AI 14.9%, Ahrefs 13.7%, Scrunch AI 9.1%. Six brands, 90% of all mentions.
Project management, 240 answers, 558 mentions, 14 distinct brands: Jira 25.3%, Asana 17.6%, Monday.com 12.4%, Linear 12.2%, ClickUp 11.1%, Wrike 4.7%. Six brands, 83%.
Project management has a bigger single leader and a longer tail. AI visibility has no runaway leader, five brands clustered between 13% and 18%, and then a cliff. Everyone outside the top six is sharing a tenth of the oxygen.
If you are selling into a young category on the theory that it is wide open, that theory is worth checking against your own measurement. Mine says the door was already closing while I was writing about how open it looked.
What I got wrong in run 1
Run 1 found eight questions where at least one engine named nobody, and I read that as eight pieces of open ground: questions being asked, no brand owning the answer, first mover takes it. I said so publicly and I built a content plan on it.
Under five repeats per engine, that mostly evaporates. Fifteen independent attempts per question later, exactly one of the sixteen questions in the young category comes back empty every single time. In the settled control there are two. A twenty-year-old market has more unowned questions than a two-year-old one, not fewer.
The eight were an artifact of asking once. An engine that names somebody two times in five will look silent if you sample it a single time, and I sampled it a single time. That is not a subtle statistical point, it is the most ordinary error there is, and the only reason I caught it is that I designed run 2 to be able to catch it.
The practical consequence is the part I care about. “Find the questions nobody owns and answer them first” is still a sound strategy. But you cannot find those questions by asking once, and a tool that asks once and shows you a gap is showing you noise about half the time. I say that as someone who sells this kind of measurement.
The one question nobody answers
The single question that came back empty on all three engines, all five times, in the young category:
How much does an AI visibility audit cost?
Zero brands named, 15 attempts out of 15.
Its exact twin in the settled category, “how much does a project management tool cost per seat,” was answered every single time, by all three engines, with named vendors and numbers.
That pairing is the cleanest result in the whole run, because everything else is held constant. Same question shape, same engines, same day. The only difference is the market. Pricing is where the young market goes unheard.
I first wrote that the reason was obvious: nobody in this category publishes prices, so there is nothing for a model to summarize. Then I went and searched, which is what I should have done before writing the sentence.
The prices are published. There are pages two and three months older than this run that answer the question directly, with dollar figures and named products. One from June names three tools and their monthly prices. One from May gives a $1,500 to $5,000 range for a done-for-you audit. An ordinary search engine surfaces both of them on the first page. Three AI engines, asked fifteen times, named none of them.
So the silence is not an empty shelf. It is three engines declining to pass on pricing information that is sitting in public, while the same three answer the identical question about project management software every single time, with vendors and numbers. I cannot prove why. The explanation I find most likely is that the pages carrying AI visibility prices are marketing pages from small vendors, and the engines weight them too low to repeat, whereas per-seat pricing for project management tools is carried by hundreds of independent reviews and comparisons that they already trust.
That is a different problem from an empty market, and it changes what you should do about it, so the advice further down reflects it rather than the version I first wrote.
I publish my own prices. That is not a coincidence and I will not pretend it is. What the audit behind that number actually measures is written up separately.
Where I place
I am a vendor in the young category, so I am in my own results. In run 1, with 47 answers, I appeared 3 times and read that as eighth place with 6% share.
With 240 answers, I appear 14 times. That is 2.1% of mentions, ninth place, behind Writesonic. The more data I collected, the worse my actual position looked, which is what usually happens when a small sample flatters you.
Publishing that is more useful to me than an estimate would be, and it is the strongest evidence I can offer that the rest of these numbers are not arranged to sell you anything.
What this study cannot tell you
Two categories is not many categories. Every effect here is measured on AI visibility tooling and project management software. A third category could break the pattern and I would have no warning.
One day, one access route. Everything was collected on 2026-08-13 through each vendor’s own interface. The St. Gallen work shows these systems move measurably inside a single day, and access route may matter more than anyone has established. My disagreement with MaxAEO may live entirely in that gap.
Sixteen question shapes, chosen by me. A different sixteen moves every number in here. They are published in full so you can judge them, and the two sets are deliberate mirror images of each other so the comparison is fair.
Being named is not being recommended. One question asks for alternatives to a named leader, and that leader appears in the answers. Presence in the text is what I counted.
Brand matching is text matching. A brand referred to in a way I did not anticipate is scored as absent.
Five repeats is enough to separate signal from noise, and not enough for confident small differences. The 30% versus 100% gap is safe to read. A five point difference between two cells is not.
The MaxAEO contradiction is unresolved. I have described it as honestly as I can and I cannot settle it from outside their data. If their methodology becomes public and my ordering is wrong, I will say so here.
Three things to do this week
If you run your own marketing and you have an hour, not a project:
Ask your five most valuable buying questions five times each, on the engine your buyers actually use. Not once. Write down only the brands that show up in at least three of the five. That list is your real competitive set, and it will be shorter and different from the one a single check gives you.
Ask what your own product costs, then check whether the answer already exists. Type “how much does a [your category] cost” into an AI engine and see whether it names a single vendor with a number. Then run the same search on Google. Those two results together tell you which problem you have. If the pages do not exist, publishing yours is cheap and you should. If the pages do exist and the engines still name nobody, the information is not missing, it is being ignored, and a seventh pricing page will not fix that. That is the situation in my category, and it took me one search to find out.
Stop aiming at “best [category] tools.” Every engine answers that one confidently, and in both markets I tested, six brands already hold more than 80% of the mentions. Displacing them is a multi-year project. Go and find the questions where the engines currently name nobody, confirm the silence by asking five times, and answer one of them properly. The longer version of how I choose which one is here.
What I am doing next
Two things, and I am naming them so you can hold me to them.
The raw data from this run is published here, both categories, all 480 answers with the brands extracted, so that anyone can check my counting or disagree with my classification:
Each record holds the question, the engine, which of the five repeats it was, and the brands named. The two academic groups in this space published their datasets. The two commercial ones did not. I would rather be in the first group.
And I am answering the pricing question in public, with my own real numbers, knowing that others have already published theirs and that the engines repeated none of them. If a page with prices on it is not enough, I would like to find out what is, and I cannot find that out without having the page.
Run 3 will repeat both categories and add a third. I am not promising a date, because run 1 promised a cadence I had not tested and that was a mistake I would rather not repeat.
If you want this run on your category instead of mine, that measurement is what I built citability.dev to do. The thing it does that a single-shot check cannot is ask more than once.
Data: 16 question shapes, 2 categories, 3 engines, 5 repeats, 480 answers, 0 failed calls, 1,230 brand mentions, collected 2026-08-13.
· Sources & further reading
Sources & Further Reading
Further reading
- I Asked 3 AI Engines the Same 16 Buying Questions. They Agreed Under a Third of the Time. /blog/ai-engine-recommendation-study A 47-answer study of who ChatGPT, Claude, and Perplexity recommend in one B2B category. Six brands hold 89% of all mentions, the three engines overlap by about 30%, and a third of answers name nobody at all.
- Why ChatGPT Is Not Citing Your Website: I Measured 47 AI Answers /blog/why-chatgpt-is-not-citing-your-website I asked ChatGPT, Claude, and Perplexity the 16 questions my buyers actually ask. My own site came up 3 times out of 47. Here is the method, the raw counts, and the gap I found.
- How to Get Cited by ChatGPT: My 0/5 GEO Audit /blog/how-to-get-cited-by-chatgpt-geo-guide I ran my own site through 5 buyer-intent queries on Perplexity and got cited zero times. Here is what GEO actually is and the pattern the winners share.
- Cloudflare Will Block AI Crawlers by Default on September 15: What Site Owners Need to Do Now /blog/cloudflare-block-ai-crawlers-september-15 Cloudflare blocks AI crawlers by default on September 15, 2026. Should you block, allow, or charge via Pay Per Crawl? Decision matrix + verification steps.
- Content Intent Signaling: The robots.txt Directive That Controls How AI Uses Your Content /blog/content-intent-signaling-robots-txt robots.txt controls access. Content Intent Signaling controls usage. Three new directives separate training from citation permission.
What do you think?
I post about this stuff on LinkedIn every day and the conversations there are great. If this post sparked a thought, I'd love to hear it.
Discuss on LinkedIn