Answer
How do I know if AI visibility work is actually working?
The short answer
To tell whether AI visibility work is doing anything, fix the question set before you start, then re-ask it. Keep the same buyer questions, the same engines and the same wording, run them on a schedule, and store every answer word for word so you can read what changed. The only honest measure of progress is how often you are named across a stable question set, not a single screenshot.
Set the baseline before you change anything
Run the full question set once, before any work starts, and keep the transcripts. A baseline is not a score. It is a list: the exact questions, the engines you asked, the date, whether you were named in each answer, and the complete text of every answer as it came back.
Without that list you cannot argue with anyone later, including yourself. With it, every claim about improvement becomes checkable arithmetic. In our own published study we asked 1,483 ChatGPT questions and 653 Gemini questions across 300 Australian businesses on 30 July 2026, and stored every answer, which is why the figures in that report can be recomputed from the raw file by someone who has never spoken to us.
Five buyer questions per engine is a workable minimum. That is the size we used per industry and city cohort in the study. Fewer than five and one lucky answer moves your whole rate.
Write the questions the way a customer would say them out loud, and do not put your business name in them. A question containing your name measures recognition, which is a different thing from being recommended. Keep the two sets separate and never average them together.
What to hold constant, and what you are allowed to change
Hold the questions, the engines and the wording constant. Your business is the variable, and it should be the only thing moving between runs.
| Element | Treatment | Why |
|---|---|---|
| The exact wording of each question | Hold constant | A reworded question is a new question. Its result is not comparable to the old one, and rewording is the easiest way to manufacture an improvement. |
| Which engines, and how they are asked | Hold constant | A ChatGPT answer with web search on is not comparable to one with it off. The same is true of Gemini with and without search grounding. |
| How many questions are in the set | Hold constant | Adding one easy question raises the rate without anything about your business improving. |
| Buyer questions and direct questions | Keep separate | They measure different things. Being recognised when asked about by name is not the same as being volunteered when a buyer asks who to use. |
| Account, chat history and location | Hold steady where you can | Worth controlling, but it matters less than the three rows above. Use a fresh chat each time. |
| Your site, listings, reviews, press and citations | This is the variable | The only thing that should be different between two runs is the world outside the measurement. |
If a vendor changes the question set mid-engagement, the history stops at that point. That is not automatically dishonest, because question sets do need to be revised, but the old and new sets cannot be graphed as one line and any report that does so is wrong.
Reading a transcript rather than a chart
Open the stored answers and read them, because the sentence an engine wrote about you carries information no score can hold. A chart tells you the count moved. The transcript tells you what actually happened.
Things a transcript shows and a number hides:
- You were named, but last in a list of seven, after six competitors.
- You were named with the wrong suburb, the wrong service or an old price.
- You were named because a directory page was cited, not because your own site was read.
- You were not named, but the answer cited three pages you could realistically get onto.
- The engine hedged the whole answer and named nobody, so the run says nothing about you either way.
That last case matters for monitoring. An answer that named nobody is not evidence that you slipped. It is a question that failed to produce a recommendation, and it should be visible as such rather than buried inside a rate.
This is also why we count the way we do. In our own engine, mentions are matched with a regular expression over the stored text and no model is ever asked to produce a number. Two inputs are model-graded from those same transcripts, a recognition flag and a sentiment value, and together they can move an engine score by at most 30 of 100 points. That split is published on the technology page rather than presented as fully mechanical.
Separating real movement from engine noise
Assume any single change of one or two mentions is noise until the same question set repeats it. These systems are not deterministic, and the pages they retrieve move underneath them, so two identical questions on the same day can return different businesses in a different order. We cover the mechanism in why AI answers change every time you ask.
Four rules that hold up in practice:
- Compare rates, never single questions. Named in 3 of 5 against 1 of 5 is a signal. Question four flipping is not.
- Require two consecutive runs. A change that appears once and is gone next month was never a change.
- Treat a one-engine move carefully. ChatGPT and Gemini disagree constantly. In our study, across the 137 businesses with complete records from both engines, Gemini scored the same business higher than ChatGPT for 105 of them, ChatGPT was higher for 7, and 25 were tied. Across all 137 the two engines' scores differed by an average of 25.3 points out of 100.han ChatGPT 77% of the time, with a mean gap of 25.3 points.
- Look for a cause in the world. If nothing changed outside the measurement, be suspicious of a change inside it.
Our own study is a snapshot, not a trend, and we say so on its own page. Each question was asked exactly once, with no repetition runs, so a business one answer short of a mention on the day stays unmentioned in that data. If you are using our figures as a benchmark, use them as a picture of one day rather than as a moving average.
What a monitoring report must contain to be worth paying for
A monitoring report earns its fee when you can check it yourself without asking the vendor a single question. Seven things make that possible.
- The exact questions asked, printed in full, not summarised as "a range of buyer queries".
- Which engines were asked, and how, including whether web search or search grounding was enabled.
- The date of each run.
- Every answer stored word for word and readable by you, not just a count.
- The result expressed as a rate with its denominator: named in 3 of 5, not "visibility improving".
- The scoring formula, published, so the number can be recomputed and argued with.
- The limits of the method, written down by the vendor rather than left for you to discover.
We hold ourselves to the last one in a specific way. Our technology page records a live defect in our own reporting: when Gemini's search grounding falls back, the run counts that fallback in its own metadata, but the report page still describes the direct verdict as grounded. On a quota-exhausted run our report therefore reads more confident than the run actually was. It is written on the method page because a limit you have to be told about by someone else was never a disclosed limit.
If you are comparing quotes, what AI visibility costs in Australia sets out the four price bands and what belongs in a fair one.
Related questions
What people ask next.
Answer
Why do I get a different answer every time?
Two sources of variation, and why a rate is the only honest unit.
Read the answer →Answer
How often should I check my AI visibility?
Monthly for most businesses, weekly while you are actively changing things.
Read the answer →Answer
How do I track ChatGPT traffic in GA4?
The referral hostnames to watch, and what the numbers will always miss.
Read the answer →Method
The Receipts Engine
How we count: the scoring formula and the limits, published in full.
Read the answer →Get the baseline before you spend anything.
The free scan asks real buyer questions of ChatGPT and Gemini, counts how often you are named, and shows you one real quote of what an engine said about you today. No email gate.