How to Track Your Brand in ChatGPT for Free, and Where It Breaks
A step-by-step manual method to see whether AI assistants name your brand, the mistake that makes most DIY tests worthless, and the point where it collapses.
You do not need to buy anything to find out whether AI assistants mention your brand. You need a spreadsheet, an afternoon, and the discipline to run the test properly. Most people who try get a useless answer, not because the method is hard, but because they make one specific mistake in the first minute.
Here is the method, in full. Then, because it matters more than the method, here is exactly where it stops being usable.
Build the prompt sheet first
Open a spreadsheet. One row per prompt. Write them before you open ChatGPT: if you improvise the questions while you type them, you will unconsciously write the ones you already know you win.
The most common failure is asking the model about yourself. "What is [your brand]?" tells you nothing: the model will produce something, it always produces something. That is a lookup, not visibility. Write the questions a buyer asks before they know you exist. They fall into three groups, and you want all three:
- Category prompts. "What is the best tool for [job to be done] for a mid-sized company?" "Who are the main suppliers of [product] in Europe?" This is the raw question of inclusion: does the model reach for your name at all when it lists the field?
- Comparative prompts. "Compare the leading options for [category]." "[Known competitor] versus alternatives." These tell you which set the model files you under, and whether you appear as a peer or an afterthought.
- Problem prompts. No category, no vendors, just the pain. "My warehouse keeps running out of stock on fast-moving items, what should I do?" Nobody tests these, and they are often where an unexpected competitor gets named, because the model is reasoning from a symptom rather than from a product taxonomy.
A dozen prompts is enough to start. Boring, specific, written in the words your customers use, not your category's internal jargon.
Use a clean session. This is the part everyone gets wrong
If you run these prompts in your normal ChatGPT account, the account that knows who you are, with your company in its memory and weeks of conversations about your own product in its history, the answer you get is not the answer a stranger gets. It is personalized, shaped by everything you have already told it. People run this test, see their brand named first, and conclude they are doing well. They are reading their own reflection.
So, before anything else:
- Turn off memory and personalization, or use a temporary chat where the assistant offers one.
- Log out entirely where you can, or use a fresh browser profile with no session.
- Never name your brand, in the prompt or in the follow-up. The moment you do, you have contaminated the rest of the conversation.
- Start a new conversation for every prompt. Not a new message in the same thread: a new thread. Context carries forward, and a model that just recommended you will keep recommending you.
This is unglamorous and it is the difference between a test and a placebo.
Run every prompt more than once
Language models are probabilistic. Ask the identical question twice and you can get two different lists of recommended vendors. This is not a defect you can tune away, and it has a consequence most manual tests refuse to accept: a single run is an anecdote. Run each prompt several times, in separate clean sessions, and treat the result as a frequency rather than a fact. "Named in four out of six runs" is a finding. "Named" is not.
Then do the same on the other assistants. ChatGPT, Claude, Gemini and Perplexity are not interchangeable: they were trained differently, they retrieve differently, and a brand confidently recommended by one can be invisible on another.
What to record for each run
Four columns. Resist the urge to add more.
- Mentioned: yes or no. Named in the answer at all.
- Linked: yes or no. Given as a clickable source or citation. Genuinely different outcomes. A model can describe your product accurately, at length, and never link to you, because it is answering from what it absorbed in training and there is no source to point at. You can be highly visible and get no traffic at all, which is why your analytics dashboard cannot see any of this.
- Who else was named. Every other brand in the answer, in the order they appear. The most uncomfortable column, and the most valuable. The companies a model names instead of you are frequently not the competitors on your internal battlecards: models group brands by how the web talks about them, not by how your market map is drawn.
- How you are described. Not a sentiment score, just the actual words. "Enterprise", "budget option", "for developers", "a smaller alternative to X". Being named with the wrong description can be worse than not being named: a model that introduces you as the cheap option has made a positioning decision on your behalf, and repeats it to every buyer who asks.
Perplexity is where you find out why
The other assistants tell you what they think. Perplexity shows you where it got it, because it retrieves live and cites its sources inline. That makes it the most useful surface in this exercise. Run your category and comparative prompts there, then ignore the answer and read the source list: those URLs are the pages the model is using to form an opinion about your market, and about you.
Expect a surprise. They are usually not your pages. They are review aggregators, a forum thread, a listicle by someone with no stake in your category, a competitor's comparison page describing you in terms you would never have chosen. That is the raw material of your AI visibility, and you have never read it.
Check that you have not locked the door
This takes a few minutes, and it is the one thing here that can invalidate everything else. Open yourdomain.com/robots.txt and look for these names:
GPTBot, OpenAI's crawlerClaudeBot, Anthropic'sPerplexityBotGoogle-Extended, which governs Google's use of your content in its generative products
A Disallow: / under any of those means you have opted that system out of reading you. Many sites added those lines during the scraping panic, on the advice of someone who has since left, and never revisited it. You may still want the block. It should be a decision, not an inheritance.
Now the honest part: where this method dies
Everything above works. It costs nothing, and you will learn more in an afternoon than from any amount of speculation about AI search. It also does not survive a second week.
Look at what you built. A dozen prompts. Each repeated several times, because one run is noise. Each of those on four assistants. And if you sell abroad, all of it again in every language you sell in, because visibility does not travel: a brand a model recommends confidently in English can be absent from the same question in German, where the web says less about it. Multiply those together. The number is not large, it is absurd. And that is one snapshot.
And here is what actually breaks the manual method: a one-off check is close to worthless. It tells you where you stood on a Tuesday. It cannot tell you whether you are drifting out of the answer set, whether a competitor is climbing, or whether the content work you spent a quarter on changed anything. The value is in the trend, and a trend means doing all of it again, identically, on a schedule, indefinitely.
Nobody sustains that by hand. That combinatorial problem is the one a monitoring tool exists to solve, and it is what PSentry automates.
Start with the spreadsheet anyway. Do it once by hand, properly, in clean sessions. You should know what is being measured before you let anything measure it for you.
Frequently Asked Questions
How many times should I repeat each prompt?
Enough that a pattern separates from noise, in a clean session each time. If your brand appears in half the runs and vanishes in the rest, that is the finding. Flattening it into a yes or no throws away the only interesting thing you learned.
Can I just use the API instead of the chat interface?
You can, and it removes the memory contamination problem cleanly. But it changes what you are testing: the model behind the API is not always configured the way the consumer product is, and may not have the same retrieval or browsing behaviour switched on. If your buyers use the chat app, test the chat app.
The model describes my product wrong. Can I correct it?
Not directly. There is no submission form and no ranking factor to tune. What a model says about you reflects what the web says about you. If Perplexity's sources for your category are a handful of review sites and a forum thread, those are the surfaces that matter, and they are not yours. Anyone offering to guarantee what an AI says about your brand is selling something they cannot deliver.
How often does a check need to be repeated to mean anything?
Often enough to see change, not so often that you are watching noise. Models do not revise their opinion of a brand from one day to the next, so daily checking mostly produces the illusion of activity. What matters is a consistent cadence, run identically each time, so that when something moves you can tell it apart from the model simply being probabilistic.