Answer
Why do I get a different answer every time I ask ChatGPT?
The short answer
Because these systems are not deterministic and the web underneath them keeps moving. The same question asked twice can return different businesses, different sources and a different order. That is why one screenshot proves nothing in either direction, and why any honest AI visibility measurement asks a fixed set of questions repeatedly and reports how often you were named rather than whether you appeared once.
Two sources of variation: the model and the web
There are two, and separating them explains almost everything you will see. The model samples its own words, and the pages it retrieves change underneath it.
The model. A large language model writes one piece of a word at a time by drawing from a probability distribution over what could come next. That draw is a sample, not a lookup, so the same input can produce different output. Both providers expose the controls in their developer documentation: OpenAI documents temperature and top_p on its API, and Google documents the same sampling parameters on the Gemini API. See the OpenAI API reference and the Gemini API documentation. The consumer chat apps do not hand you those dials, but the sampling is still happening.
The web. ChatGPT with web search, and Gemini with Google Search grounding, fetch live pages at the moment you ask. Which pages come back depends on what is currently indexed and currently ranking, and that moves hour to hour. A competitor publishing a comparison page, a directory reshuffling, a review landing, all of it changes the material the answer is written from.
There is a third, smaller source worth naming so you do not mistake it for the first two: your own account. Chat history, saved memory and approximate location can all shape a reply. Ask in a fresh chat with memory off if you want a cleaner read.
What this means for a screenshot a salesperson shows you
A single screenshot proves almost nothing, in either direction. It is one sample from a distribution, taken on one account, on one day.
Applied honestly, that cuts both ways:
- A screenshot showing your competitor named and you absent does not establish that you are invisible.
- A screenshot showing you named does not establish that you are visible.
- A before and after pair of screenshots is the weakest evidence in this category, because the difference between them can be produced by asking twice.
You can test the claim in about a minute. Open a fresh chat, ask the same buyer question three times in three separate chats, and compare. If the three answers name different businesses, the screenshot method is dead and anybody selling on it knows or should know that. Our own view on the evidence in this market, from a firm that sells the work, is on is AI visibility worth it.
Why a rate is the only honest unit
The honest unit is a rate: how many questions from a fixed set named you, out of how many were asked. Named in 3 of 5 is a fact somebody else can reproduce. "Visible on ChatGPT" is not.
This is why every figure in our own research carries a denominator. Across the 300 businesses we scanned on 30 July 2026 we asked 1,483 ChatGPT questions and 653 Gemini questions, and the published findings are stated as rates over an explicit sample:
| Finding | Figure | Counted over |
|---|---|---|
| Never named once by ChatGPT | 54% | All 300 businesses (162 of 300) |
| Never named once by Gemini | 48% | The 137 with a complete Gemini record (66 of 137) |
| Never named by either engine | 41% | The 137 with complete two-engine records (56 of 137) |
| Recognised when asked directly, never named when a buyer asks | 26% | The same 137 (35 of 137) |
A number with no denominator is not a measurement, it is a mood. If a report gives you a percentage without telling you what it is a percentage of, that is the first thing to ask about. The complete set of our figures, each with its denominator, is on the statistics page.
How many times you need to ask before a change is real
More than once, and we are not going to give you a threshold we have not measured. We have run no repetition study, so any specific number from us would be invented, and invented numbers are the problem this page exists to describe.
What the structure of the problem does support:
- Repeat the whole set, not one question. A rate across five or more questions is far steadier than any single question, because one flipping answer moves a five-question rate by 20 points and a twenty-question rate by 5.
- Require the change to survive a second consecutive run. A move that appears once and vanishes was never a move.
- Watch both engines. A change in one and not the other is weaker evidence than a change in both. In our sample, Gemini scored the same business higher than ChatGPT 77% of the time across the 137 with complete records from both engines, so the two are not interchangeable readings of the same thing.
- Ask what changed in the world. A change in the measurement with no corresponding change outside it deserves suspicion, not a slide.
The practical cadence that follows from this is covered in how often you should check your AI visibility.
What we do about it in our own scans
We ask a fixed question set, store every answer word for word, and count mentions with a regular expression rather than asking a model for a number. The design is a direct response to the variation described above.
- The per-engine score is a rate, mentions divided by questions asked, not a yes or no.
- Matching is deterministic. A model proposes the brand names a business trades under; those names become the expressions, and the matching that produces every count is machine text matching. Short names of six characters or fewer with no space only match when capitalised, because short brand names are usually ordinary English words too.
- 70% of an engine score is that count. The remaining 30% is model-graded recognition and sentiment, read from the same stored transcripts, and we say so rather than presenting the number as fully mechanical.
- The stated limits sit on the same page: two engines rather than five, scheduled rather than realtime, no access to anyone's real ChatGPT sessions, and each question in the published study asked exactly once.
All of it is on the technology page, including the parts that do not flatter us.
Related questions
What people ask next.
Answer
How often should I check my AI visibility?
Monthly for most businesses, weekly while you are actively changing things.
Read the answer →Answer
How do I know if the work is actually working?
Fix the question set, then re-ask it. Read transcripts, not screenshots.
Read the answer →Answer
How do I check what ChatGPT says about me?
The two questions to ask, in order, and how to read the replies.
Read the answer →Method
The Receipts Engine
How we count: the scoring formula and the limits, published in full.
Read the answer →Ask it properly, once.
The free scan puts a fixed set of real buyer questions to ChatGPT and Gemini, counts how often you are named, and shows you one real quote of what came back. No email gate.