Back to blog The Same Technical Question, Four Different AI Answers

The Same Technical Question, Four Different AI Answers

Ask four assistants who can machine a part to your spec and you get four different supplier lists. The reasons are structural, and testable in an afternoon.

Take a question a buyer would actually type: who can machine titanium parts on a five-axis centre, in Europe, with aerospace quality certification. Paste it word for word into ChatGPT, Claude, Gemini and Perplexity on the same afternoon. You will get four different supplier lists. Some names will overlap, some will appear only once, and at least one assistant will name a distributor where another names a manufacturer. Nothing has gone wrong. Four different systems answered four times, and the differences between them are structural rather than accidental.

An assistant is a model plus everything around it

It is tempting to treat the four as interchangeable windows onto the same knowledge. They are not. Each one is a language model wrapped in machinery: a decision about whether to search the live web, a set of instructions telling the model how to behave, a cutoff date on what it absorbed during training, and a sampling process that introduces variation even when everything else is held still. Change any one of those layers and the supplier list changes with it.

For a company selling technical products, this matters in a specific way. You are not trying to rank in one place. You are trying to be findable by four systems that disagree with each other about which sources are worth reading.

What actually differs

What the model absorbed, and when it stopped

Every model has a training cutoff, a point after which it learned nothing new. The cutoffs are not aligned across providers, and neither are the corpora behind them. A product line you launched last quarter may sit inside one model's training window and outside another's. The same is true of a rebrand, a factory acquisition, or a certification you obtained recently. When an assistant answers purely from what it absorbed, it is describing a version of your market that stopped at a particular moment.

Whether it looks anything up before answering

This is the largest single source of divergence. Some assistants are built to search first and compose an answer from what they just fetched. Others answer from what the model already contains, and search only when the question seems to require it, or when you ask them to. The first behaviour rewards pages that are fetchable, current and quotable. The second rewards being written about widely enough, and for long enough, that the model absorbed your name in the first place.

These are two different games. A supplier with an excellent, freshly updated product page and almost no third-party coverage tends to do better with search-first assistants. A long-established name that appears in industry directories, association listings and trade publications tends to survive better in the answers that come from memory alone.

How the answer is sampled

Language models are probabilistic. Given the same prompt twice, they can produce different wording and, more importantly for you, different names in different order. That is not a defect to be fixed, it is how the generation works. It also means a single test tells you very little. Asking once and concluding that an assistant does not know you is like calling a warehouse once, getting a voicemail, and concluding the warehouse is empty.

What the surrounding instructions tell it to do

Each product wraps its model in instructions about tone, caution, citation and refusal. One will hedge and refuse to recommend named vendors for a purchasing decision. Another will produce a numbered shortlist without hesitation. A third will name companies but attach a caveat about verifying certifications directly. This is a product design choice, and it changes what a buyer sees even when the underlying knowledge is similar.

Run the comparison properly

The test is worth doing yourself before you accept anyone's summary of it, including ours. It needs an hour and no tooling.

  • Write the query the way a buyer would. Not your brand name. A technical need: process, material, tolerance class, certification, region.
  • Use identical wording in all four. Any rephrasing invalidates the comparison, because the phrasing is part of what is being tested.
  • Log three things per answer: which companies are named, in what order, and which sources the assistant cites if it cites any.
  • Repeat the whole set the next morning. What changes between days is variation. What stays put is signal.
  • Then run it again in another language you sell in. This is where most exporters find the widest gap, and it deserves its own session rather than a footnote to this one.

What to do with four disagreeing answers

The instinct is to pick the assistant your buyers use most and optimise for it. That is the wrong read, mostly because you cannot know which one a given buyer used, and partly because optimising for one system's retrieval quirks is a bet on machinery that changes without notice.

The useful reading is the pattern across all four. A company named by every assistant is present in a structural way: in enough independent places that both memory and live retrieval find it. A company named by exactly one, especially by a search-first assistant, often owes that mention to a single page that happened to rank well that day, which is a thinner position than it looks. And when three assistants name your distributor rather than you, the problem is usually not the assistants: it is that the distributor's page states the specifications in plain text while yours hides them in a downloadable file.

The sources cited are worth more than the ranking. They tell you which pages the machinery considers authoritative for your category, and that list is often uncomfortable reading: a marketplace listing, a directory entry, a competitor's comparison page, a forum thread. Those are the surfaces where your category is being described, whether or not you participate in them.

The boundary of this exercise

These four assistants are not the whole of AI search. Google's overview features and the assistants embedded in office software are separate surfaces with their own behaviour. We do not cover them, and it is worth being explicit about that rather than implying a completeness we do not have. PSentry measures how your brand appears across ChatGPT, Claude, Gemini and Perplexity, per language and per market, on a scheduled basis rather than continuously. It measures. It does not influence what any assistant says about you, and no monitoring tool can promise that it will.

What a structured measurement adds over the manual test is repetition and coverage: the same prompts, across the four systems and every language you sell in, sampled over time so that the difference between noise and a real change becomes visible. The manual test tells you whether the problem exists. Repetition tells you whether it is moving.

Frequently Asked Questions

If the four disagree, which one is correct?

None of them is authoritative. They are different systems reading different sources at different moments, so disagreement is the expected output, not an error to be resolved. Treat the set of answers as a sample of how your category is described, not as a ranking to be won.

Why does the same assistant answer differently when I ask twice?

Generation is probabilistic, and any assistant that searches the live web is also reading a web that changed between your two attempts. Both effects push in the same direction: a single answer is a snapshot, and conclusions need repetition.

Does asking in German or Spanish change the supplier list?

Frequently, and more than most exporters expect. The available sources differ by language, and so does the volume of material a model absorbed in each. A brand that appears reliably in English answers can be missing entirely from the same question asked in another language.

Should I write separate content for each assistant?

No. Retrieval behaviour changes without notice, and content built around one system's current quirks ages badly. What survives across all four is the same unglamorous thing: your technical facts stated as ordinary text, on a page that can be fetched, in each language you sell in.

Do the assistants use the same model in every country?

Not necessarily, and the wrapping around the model can differ too, including which sources are reachable. This is another reason to run your comparison in each market you care about rather than assuming the English result generalises.