Why does AI give a different answer every time, and what does this mean for measurement?
AI responds differently each time because it generates text with an element of randomness; each time it is run, it may search different sources, and the result also depends on the account, location and model version. The variation is significant: in a study by SparkToro and Gumshoe.ai, ChatGPT and Google’s AI returned the same list of brands in fewer than 1 in 100 repetitions of the same question. What is more consistent is how often a brand appears across multiple responses, and it is precisely this metric that needs to be measured.

This text is part of the guide Brand visibility in AI: how it works and how to measure it.
How much do the answers to the same question differ?
More than most marketers realise. A study by SparkToro and Gumshoe.ai From the start of 2026, the study involved 600 volunteers, 12 questions and 2,961 queries in ChatGPT, Claude and Google’s AI search function. The questions concerned, amongst other things, kitchen knives, headphones, cancer hospitals and marketing consultants (a feature on Search Engine Land).
Results:
- the chance that ChatGPT or Google’s AI would provide the same list of brands in two different replies was less than 1 in 100,
- the same list in the same order appeared roughly once in every 1,000 attempts,
- The length of the list also varied: sometimes the model would list 2–3 brands, and just as often 10 or more,
- Claude tended to repeat the same set of brands slightly more often, but even less frequently in the same order.
- 1 in 100the same list of brands
- 1 in 1,000the same list in the same order
- 2 to 10+So many brands give one answer one time and another answer the next
Where does this volatility come from?
Randomisation during text generation
The model constructs its response word by word, selecting each word from a probability distribution. In chat conversations, this selection process includes an element of randomness, which makes the text sound natural. A side effect of this is that, when two brands are equally likely, one wins one time and the other the next.
Other searches every time the programme is launched
When the chatbot uses the internet, it breaks the question down into sub-queries. The next time it runs, it may generate different queries, retrieve different web pages and cite different sources. In the AirOps study ChatGPT cited only 15% pages that it had retrieved, so the very selection of sources for the response is itself a form of selection. The mechanism is described in the text How AI chatbots choose the brands they recommend.
Different functions of the same platform
AI Mode and AI Overviews on Google are two different products. Google points out in Search Central documentation, …that they may use different models and techniques, so they display different results and links. According to Ahrefs data from December 2025, for the same query, they cite different sources in 87% cases (a feature in Search Engine Journal).
Account, storage and location
A logged-in user with a chat history, ChatGPT’s memory and, based on their set location, receives a different response to someone in temporary chat. When it comes to local services, the difference can be significant, which is why companies operating in a single town or city should check for queries containing the name of the locality (more on the website Aelo for small businesses).
Updates to models and indices
Providers change their models and search methods without warning, and website indexes are updated daily. Lily Ray describes, how ChatGPT’s supporting queries and the sources it uses have changed in 2026. A result from a month ago may look different today, even though nothing has changed on your website.
What remains constant in AI responses?
The pool of brands from which consumers choose. In the SparkToro survey, the leading brands in the headphones category appeared in 55–77% of the responses, although the order and composition of the lists varied almost every time. In narrow categories, such as cloud computing providers, the biggest players featured in the majority of responses. In broad categories, such as science fiction novels, the results were much more scattered.
The questions themselves also vary. When 142 survey participants wrote their own questions about headphones, their semantic similarity was just 0.081. Customers ask about the same thing in very different ways.
What does this mean for measuring brand visibility?
Since the AI gives a different response every time, the position on the list is not a reliable indicator. A brand that is „third in ChatGPT” may be seventh the next time you run it, or it may disappear altogether. A meaningful measurement is based on:
- the percentage of responses in which the brand appears,
- several versions of each question, as customers phrase them in different ways,
- multiple runs across several models, rather than a single attempt in a single chat,
- a fixed set of queries and a trend compared on a weekly or monthly basis.
- 3rd place
- 7th place
- not on the list
A screenshot from a client or boss with the caption „we’re not on ChatGPT” is just one example. It proves neither that there is a problem nor that there isn’t one. How to carry out a manual test that provides meaningful results, as described in the text How to check whether AI recommends your brand. The following illustrates what a repeated measurement on the same set of queries looks like: case of a lime manufacturer.
Is it possible to disable the random nature of the answers?
The app user has no control over this. In the API for some models, the temperature parameter can be lowered, which reduces text variability but does not eliminate differences arising from the search. The response from the API may also differ from what the customer sees in the app. When assessing visibility, what matters is what users actually see, as verified through numerous tests.
How can visibility be measured despite fluctuations?
Regularly and using the same queries. At Aelo, we do this in the module Visibility: we run a set of queries across four models (ChatGPT, Claude, Gemini, Perplexity) at regular intervals and show how the brand’s presence changes over time and relative to the competition. The results are fed into the knowledge graph Synapsis, and on this basis we draw up specific measures in Tickets. It’s up to you to decide which parts of this you’ll put into practice.
Agencies that need to explain to their clients why a single day’s result is not representative will find more information on the website Aelo for agencies. The article explains how this measurement differs from Google’s location tracking Visibility in AI and SEO.
You can start tracking with a free account. Sign up
FAQ
No. In a study by SparkToro and Gumshoe.ai, ChatGPT provided the same list of brands in fewer than 1 in 100 attempts. Even when the question is identical, the composition of the list, the order of the brands and the length of the list vary.
In the SparkToro study, each query was run 60–100 times per platform. In a manual test, 5–10 repetitions for each model are enough to see whether the brand appears regularly, occasionally or not at all.
Yes, if you are measuring the percentage of responses mentioning a brand against a fixed set of queries, rather than its position. The pool of brands from which the model selects is stable: the leading brands in the category under study appeared in 55–77% responses.
Variability refers to fluctuations between runs of the same query when nothing has changed. A real change is a shift in the percentage of responses mentioning the brand across successive measurements using the same set of queries. If the set of queries changes, results from different days cannot be compared.
Sources
- SparkToro and Gumshoe.ai, a study on the consistency of brand recommendations in AI (January 2026): https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/
- Search Engine Land, an overview of the SparkToro study (January 2026): https://searchengineland.com/ai-recommendation-lists-rarely-repeat-study-468076
- Search Engine Journal, an analysis of SparkToro’s research and Ahrefs’ data on AI Mode and AI Overviews (January 2026): https://www.searchenginejournal.com/ai-recommendations-change-with-nearly-every-query-sparktoro/566242/
- AirOps, report on searches, follow-up queries and citations in ChatGPT (March 2026): https://www.airops.com/report/influence-of-retrieval-fanout-and-google-serps-in-chatgpt
- Google Search Central, AI features and your website: https://developers.google.com/search/docs/appearance/ai-features
- OpenAI, ChatGPT’s memory: https://help.openai.com/en/articles/8590148-memory-faq
- OpenAI, temporary chat: https://help.openai.com/en/articles/8914046-temporary-chat-faq


