How to check whether AI recommends your brand
A manual brand visibility test in AI involves asking the same customer questions, without mentioning the company name, in several chat sessions and repeating them 5–10 times. A single attempt is not enough: in SparkToro’s study, ChatGPT and Google’s AI returned the same list of brands less than once in every 100 attempts. Instructions with sample questions and a measurement sheet.

- How can you prepare a test so that the result isn’t skewed?
- What questions should you ask?
- Which AI chatbots should I try out?
- What should you note down after each answer?
- How many times should each question be repeated?
- What should we do with the results?
- When is a manual test no longer sufficient?
- Five hundred answers on a sheet of paper or a single scan
- Questions
- Sources
To check whether AI recommends your brand, ask a few AI chatbots the questions that customers ask before making a purchase, without mentioning your company’s name. Repeat each question several times and note down whether the brand is mentioned, alongside which competitors, and how the model describes it. A single trial is not very telling, as the answers vary between runs. Only the percentage of responses mentioning the brand across multiple trials provides a reliable picture.
This text is part of the guide Brand visibility in AI: how it works and how to measure it. A shorter version of the instructions, without the worksheet and calculations, is available on the website How to check whether AI mentions the brand.
How can you prepare a test so that the result isn’t skewed?
The chatbot you’ve been talking to about your business for months knows it better than your client’s chatbot. ChatGPT uses memory and chat history, so they’ll be more likely to mention the brand they’ve been messaging you about. Before the test:
- use the temporary chat feature, which By default, it does not use custom memory or instructions, or disable ChatGPT’s memory in the Settings > Personalisation > Memory menu,
- Do not mention the brand name in your question or in previous messages in the same conversation; in other words, please neutral scan,
- If you’re operating locally, include the town in your query, as the results for local services depend on your location,
- Make a note of whether the chatbot used web search. Without it, you are testing knowledge from the model’s training; with it, you are testing what the model finds today. The text describes the difference between these sources How AI chatbots choose the brands they recommend.
What questions should you ask?
The sort of questions a client asks before choosing a firm. An example for an accountancy firm in Katowice:
Test kit
An example for an accountancy firm in Katowice
-
Category
„Which accountancy firm in Katowice provides services to limited liability companies?”
It checks whether the brand is even included in the pool of companies associated with the category.
-
Problem
„Who in Katowice can help a small company prepare for KSeF?”
It checks whether the brand appears in connection with a specific need, rather than simply in connection with the name of the sector.
-
Comparison
„An accountancy firm or a full-time accountant in a company with a staff of eight?”
It checks for presence at the stage where the customer is wavering between two options.
-
The brand, plain and simple
„What does this company do, and how much do its services cost?”
It does not measure visibility. It checks recognition and whether the model is providing accurate information.
Write down each question in two or three versions, and then do not change the set between measurements.
aeloapp.ioA question about a brand does not directly measure visibility. It checks diagnosis and recommendation, i.e. whether the model is even familiar with the company and whether they provide accurate details about it: current address, range of services, prices. For the other questions, stick to search intent and the stage of the customer’s decision-making process, rather than the search volume of the phrase.
Write down each question in two or three versions. In SparkToro survey 142 people asked their own questions about headphones, and the wording of their questions hardly overlapped at all (semantic similarity of 0.081). Customers ask the same thing in many different ways, and the model may suggest different brands for each of them.
Once the set-up is ready, do not alter it between measurements. We have our own name for this rule: frozen query set. Without it, the two measurements are no longer comparable.
Which AI chatbots should I try out?
The ones your customers use. These will most often be ChatGPT, Gemini, Perplexity, Claude, Microsoft Copilot and AI-generated answers in Google Search. Check Google’s AI Mode and AI Overviews separately. Google admits in Search Central documentation, that both functions may use different models, and according to Ahrefs’ data, for the same query, they cite different sources in 87% cases (a feature in Search Engine Journal).
You can check Copilot to some extent without carrying out manual tests. If the website is verified in Bing Webmaster Tools, AI Performance Report shows which Copilot subpages and AI responses in Bing are cited, and for which queries. There is no such report for the other models.
What should you note down after each answer?
The simplest way is to keep a spreadsheet with one row per answer:
Measurement sheet
One line per reply, eleven columns
| Date | test day |
|---|---|
| Model and mode | e.g. ChatGPT with search, Gemini without search |
| Question | exact wording, including the variant number |
| The brand’s response | yes or no, i.e. a mention |
| Link to your website | Yes or no: a note on citation |
| Item | position in the list and the length of the list, e.g. 3 out of 7 |
| Competitors | brands mentioned in the same answer |
| Brand description | how the model represents the company, in one sentence |
| Alternative name | under which entry did the model list the company? |
| Errors | out-of-date or false information |
| Sources | pages linked to by the model |
10 questions × 2 options × 5 models × 5 repetitions = 500 answers
to be read and transcribed onto the sheet during a single measurement
The poem „Sources” has been highlighted, as it is the most important column on the page.
aeloapp.ioThe definitions of two columns from the worksheet are in the dictionary: reference i citation. It’s also worth adding sentiment in the brand description and brand aliases for the name option.
The ‘Sources’ column is the most valuable. It shows which websites the model uses to gather information about your industry, and this provides you with a ready-made list of sites to explore.
How many times should each question be repeated?
Certainly more than once. In the SparkToro and Gumshoe.ai study, each query was run 60–100 times per platform, whilst ChatGPT and Google’s AI returned the same list of brands in fewer than 1 in 100 attempts. The article explains the reason for this variability Why does AI give a different answer every time?.
In a manual test, that many repetitions are unrealistic. However, running each question 5–10 times across each model shows whether the brand appears regularly, occasionally or not at all. The result is calculated as a percentage: a brand appearing in 6 out of 10 answers is 60%. The brand’s share compared to other companies in the same category is calculated as share of voice, and the percentage of replies containing a link to your website as citation index. Don’t treat the position on the list as an indicator, as it changes every time you run the programme.
To identify a trend, you need to repeat the measurement. We explain how often on the page How often should visibility be measured?.
What should we do with the results?
If the brand isn’t mentioned in the responses, have a look at the sources the model cites in relation to these questions. This is a ready-made list of places where your competitors are present but you aren’t: catalogues, comparison sites, trade media and review sites. We explain why the model refers to these sources rather than your website on our website Why does AI recommend the competition?.
If the model doesn't associate the company with anything at all, start with the page ChatGPT doesn’t know my company. If he recognises it but never recommends it, that’s a different issue, described in the text He knows it, but doesn’t recommend it.
If the make is correct but the model is described incorrectly, correct the information at the source: on the website, in Company Profile on Google, in the directories. The model repeats what it reads. The facts we have gathered about the brand from your website, scans and citations can be seen in Brand awareness, and the information you provide about the company is supplemented in Brand identities.
If a brand appears infrequently, compare the question variants in which it appears with those in which it does not. The difference usually indicates content gap: a missing price list, a description of the service for a specific customer group, or a comparison with an alternative. Local businesses can find more guidance on the website Aelo for small businesses, and personal brands on the website Aelo for personal brands.
When is a manual test no longer sufficient?
When you want to track more than just a few queries, compare results over time or report them to clients. The module Visibility I regularly ask the same questions across four models (ChatGPT, Claude, Gemini and Perplexity) and record the results, so you can see a trend rather than just individual screenshots. The scope of a single measurement is described by the keyword full scan, and the method for calculating the results methodology. Based on the results, we prepare tasks, e.g. a list of platforms where it is worth setting up a profile. Agencies conducting measurement on behalf of multiple clients will find further details on the website Aelo for agencies, and the terms and limits on enquiries in price list.
This shows how such a measurement changes over the course of a week case of a lime manufacturer: the same set of eighteen queries, two measurements, a five-week gap.
Five hundred answers on a sheet of paper or a single scan
The manual test described in this text works, but it takes a few hours each time it’s repeated. We run the same set of queries across four models automatically and record the results so that they can be compared with the next measurement. The Free account gives you 250 AeloCoins to start with and a full brand report.
You can find the number of queries monitored under each plan in price list.
Questions
To get a general idea, 5–10 questions, each with two options, are sufficient. Consistency is more important than the number of questions: you need to use the same set of questions for the next assessment, otherwise you won’t be able to compare the results.
You can, but the result will only apply to that particular model. Everyone uses different sources, so a brand may appear in one but not in another. A measurement in a single chat is just a single sample, not a diagnosis.
Because then you’re suggesting the answer to the model. A question that includes the name checks whether the model recognises the company and describes it correctly. A question without the name checks whether the model mentions it of its own accord, and that alone is the measure of visibility.
Correct them at source: on your own website, in your business card and in directories, then repeat the measurement after a few weeks. The model repeats what it reads, and changes to the sources take time to take effect, as we explain on the page Why do changes take time?.
Once a month, on the same day and using the same set of queries. More frequent measurements mainly reveal the variability of the models, rather than the results of your work. We explore this further on the page How often should visibility be measured?.
Sources
- SparkToro and Gumshoe.ai, a study into the consistency of brand recommendations in AI (January 2026): https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/
- Search Engine Journal, a review of SparkToro’s research and Ahrefs’ data on AI mode and AI Overviews (January 2026): https://www.searchenginejournal.com/ai-recommendations-change-with-nearly-every-query-sparktoro/566242/
- OpenAI, temporary chat: https://help.openai.com/en/articles/8914046-temporary-chat-faq
- OpenAI, ChatGPT’s memory: https://help.openai.com/en/articles/8590148-memory-faq
- Google Search Central, AI features and your website: https://developers.google.com/search/docs/appearance/ai-features
- Search Engine Land, ‘AI Performance’ report in Bing Webmaster Tools (February 2026): https://searchengineland.com/bing-webmaster-tools-ai-performance-report-468751


