Back to blog Why ChatGPT, Claude, Gemini and Perplexity Disagree About You

Why ChatGPT, Claude, Gemini and Perplexity Disagree About You

Ask the four big AI assistants the same buying question and you often get four different brand lists. Watching only one of them measures a quarter of your visibility.

Take a question one of your buyers would actually type. Not your brand name: a real buying question, the best tool for X, the main suppliers of Y in Europe, who to consider for Z. Paste it into ChatGPT. Then Claude. Then Gemini. Then Perplexity. Same wording, same day, four tabs.

Read the four answers side by side and they are not variations of one another. They are different lists. Some brands appear in three and vanish in the fourth. Some appear in exactly one. Occasionally the lists barely intersect at all.

This breaks the mental model people bring from search. With Google there was one index and one set of results, and everyone was arguing about the same scoreboard. Here there is no shared scoreboard, only four systems that learned about your market in different ways.

Measuring one platform is not a cheaper version of measuring all of them

It is a different measurement, and it can be wrong in a direction you will not detect. Check ChatGPT and find yourself present, and you have learned that ChatGPT names you. About the other three you have learned nothing. Not "probably fine", not "roughly similar". Nothing. Correlation between platforms is not yours to assume, because the mechanisms deciding who gets named are not the same mechanisms.

Where the divergence actually comes from

Live retrieval versus what the model already absorbed

These systems can answer in two very different ways. They can fetch pages from the web at the moment you ask and build the answer out of them. Or they can answer from what they internalised during training, with no page fetched at all. Most can do both, and which one happens depends on the product mode, the phrasing of the question, and vendor decisions you do not get to see.

Perplexity is the clearest case: it is built around retrieval and shows the sources it used as a matter of course. The others sit on a spectrum that shifts with the mode and the question. We will not pretend to know the exact rules: the vendors do not publish them, and they change.

The consequence is concrete. A brand the live web talks about a lot, but that no model has strongly internalised, shows up well in retrieval-heavy answers and thinly elsewhere. A brand discussed for years that has gone quiet can show the opposite. Same brand, same question, two verdicts, for a structural reason that has nothing to do with your marketing this quarter.

Citation is not a courtesy, it is a consequence

People read a missing link as the model hiding its homework. It usually is not. A model answering from training has no source to point at, because there is no document in front of it. It is producing text from parameters, not summarising a page it just read. There is nothing to cite.

So citations are not a UX preference. They are a readout of how that answer was produced. Which is why being mentioned and being cited have to be counted as two separate events: you can be described accurately, at length, by a system that will never send you a single visit.

Each platform reads the web with its own bot, and you may have blocked one

This is the most checkable item on the list, and the one that most often turns out to be quietly broken.

The platforms do not share a crawler. OpenAI publishes GPTBot, and separately OAI-SearchBot for its search surface. Anthropic publishes ClaudeBot. Perplexity publishes PerplexityBot. Google is the odd one out: Google-Extended is not a crawler but a control token, a way of telling Google whether content it already fetched with Googlebot may be used for its generative products. Each vendor documents its own current list, and those lists have changed before and will change again, so read the vendor's page rather than a blog post, this one included.

Now open yourdomain.com/robots.txt and look. A Disallow: / under one of those names and not the others is an asymmetry you built into your own visibility. Plenty of sites added a rule for one bot during a scraping panic, or inherited one from a security plugin, a CDN default, an agency that no longer works there. Block one of the four and you should expect one of the four to know less about you. That is not a mystery of the algorithm. It is a line in a text file.

The models move

These systems get retrained, updated and swapped underneath the same product name. What a platform says about your category today is a photograph, not a law. That is not a reason to stop measuring. It is a reason to measure repeatedly, and to treat any single reading on any single platform as a data point rather than a finding.

Visibility does not transfer

The operational conclusion is blunt: strength on one platform tells you nothing about the others, and there is no channel through which it propagates. No central registry of brands exists that all four consult. Being one system's favourite does not put you in another's training data, does not get you fetched by another's crawler, does not make another's retrieval surface your page.

So the gap between platforms is not noise to be averaged away. It is itself information, and it is diagnostic:

  • Present only where live retrieval dominates. The web says useful things about you, but the models have not internalised you. Your presence depends on a page being fetched in real time. Real, and fragile.
  • Present in training-heavy answers, absent in retrieval-driven ones. You had a reputation and the current web is not reinforcing it. Or your pages are not being fetched, which sends you back to the robots.txt paragraph above.
  • Present in three, absent in one. Look for a mechanical cause before a strategic one. A blocked bot, or a market where that platform has thin material to work with.

None of that reads as a single score. It reads as a shape, and you only see the shape if you are looking at all four.

And now multiply it by every language

Everything above happens independently in each language you sell in. The four platforms disagree in English, they disagree differently in German, and the German disagreement is not a translation of the English one. What the models absorbed about your category from German-language sources is a different body of material with different brands in it, and what gets retrieved live in German comes from different pages.

So the matrix is platforms times languages, and every cell can hold a different answer. A brand can be strong on two platforms at home and absent from all four in an export market where it has real revenue. Nobody finds that by checking one assistant in one language.

What that means for how you measure

Run the four-tab test by hand today. It is the fastest way to stop arguing about whether any of this matters. What you cannot do by hand is run it across a realistic prompt set, on four platforms, in every language you sell in, often enough that the probabilistic noise cancels out. That is arithmetic, not ambition.

That is the point at which an AI monitoring tool stops being a luxury, and it is what PSentry does: it runs your prompts across ChatGPT, Claude, Gemini and Perplexity, in each language and market, and reports where you are named, where you are cited, which sources were used, and who was named instead of you.

The honest limits, since the whole point here is that unexamined assumptions cost you: a photograph across all four platforms, taken twice inside a month, is what you get, not a live feed and not a lever, and nothing about it comes with a promised result. What you get is an accurate picture of what they say now, on all four, instead of an extrapolation from the one you happened to check.

The uncomfortable part

Most brands that "monitor AI" are monitoring ChatGPT, because it is the one everyone has heard of, and reporting the result as though it described their position in AI generally. It describes their position in ChatGPT. The other three are a separate question, and until you ask it, the correct thing to say about your visibility there is that you do not know.

Frequently Asked Questions

If the four platforms disagree, which one should I care about most?

The one your buyers use, which your analytics probably cannot tell you. Absent that, treat them as four separate markets rather than ranking them. Watching all four costs less than confidently reporting on one and being wrong about the rest.

Does blocking a crawler mean I disappear from that platform?

Not necessarily, and the distinction matters. Blocking a bot limits what that platform can fetch from your site. It does not remove what the platform already absorbed, and it does not stop it learning about you from what other people publish. It removes your ability to speak for yourself, which is not the same as being erased.

Can I be visible on Perplexity and invisible on Claude at the same time?

Yes, and it is a common, readable pattern. It usually means the live web has material about you that a retrieval-driven system can find, while the model itself has not internalised your brand as an answer to that question. Presence built on retrieval alone is real but thin.

Do I need to write different content for each platform?

No, and anyone selling you platform-specific content is selling something they cannot substantiate. There is no submission form and no per-platform ranking factor. What differs is how each system reaches the web, not what it wants to find there. Clear, fetchable pages and other people writing about you are the whole of what you control.