Answer

How do I know if AI visibility work is actually working?

The short answer

To tell whether AI visibility work is doing anything, fix the question set before you start, then re-ask it. Keep the same buyer questions, the same engines and the same wording, run them on a schedule, and store every answer word for word so you can read what changed. The only honest measure of progress is how often you are named across a stable question set, not a single screenshot.

Answered 1 August 2026Method: The Receipts EngineAnswerable method, not a measured finding

Set the baseline before you change anything

Run the full question set once, before any work starts, and keep the transcripts. A baseline is not a score. It is a list: the exact questions, the engines you asked, the date, whether you were named in each answer, and the complete text of every answer as it came back.

Without that list you cannot argue with anyone later, including yourself. With it, every claim about improvement becomes checkable arithmetic. In our own published study we asked 1,483 ChatGPT questions and 653 Gemini questions across 300 Australian businesses on 30 July 2026, and stored every answer, which is why the figures in that report can be recomputed from the raw file by someone who has never spoken to us.

Five buyer questions per engine is a workable minimum. That is the size we used per industry and city cohort in the study. Fewer than five and one lucky answer moves your whole rate.

Write the questions the way a customer would say them out loud, and do not put your business name in them. A question containing your name measures recognition, which is a different thing from being recommended. Keep the two sets separate and never average them together.

What to hold constant, and what you are allowed to change

Hold the questions, the engines and the wording constant. Your business is the variable, and it should be the only thing moving between runs.

What a comparable re-scan holds constant, and what it is allowed to vary
ElementTreatmentWhy
The exact wording of each questionHold constantA reworded question is a new question. Its result is not comparable to the old one, and rewording is the easiest way to manufacture an improvement.
Which engines, and how they are askedHold constantA ChatGPT answer with web search on is not comparable to one with it off. The same is true of Gemini with and without search grounding.
How many questions are in the setHold constantAdding one easy question raises the rate without anything about your business improving.
Buyer questions and direct questionsKeep separateThey measure different things. Being recognised when asked about by name is not the same as being volunteered when a buyer asks who to use.
Account, chat history and locationHold steady where you canWorth controlling, but it matters less than the three rows above. Use a fresh chat each time.
Your site, listings, reviews, press and citationsThis is the variableThe only thing that should be different between two runs is the world outside the measurement.

If a vendor changes the question set mid-engagement, the history stops at that point. That is not automatically dishonest, because question sets do need to be revised, but the old and new sets cannot be graphed as one line and any report that does so is wrong.

Reading a transcript rather than a chart

Open the stored answers and read them, because the sentence an engine wrote about you carries information no score can hold. A chart tells you the count moved. The transcript tells you what actually happened.

Things a transcript shows and a number hides:

  • You were named, but last in a list of seven, after six competitors.
  • You were named with the wrong suburb, the wrong service or an old price.
  • You were named because a directory page was cited, not because your own site was read.
  • You were not named, but the answer cited three pages you could realistically get onto.
  • The engine hedged the whole answer and named nobody, so the run says nothing about you either way.

That last case matters for monitoring. An answer that named nobody is not evidence that you slipped. It is a question that failed to produce a recommendation, and it should be visible as such rather than buried inside a rate.

This is also why we count the way we do. In our own engine, mentions are matched with a regular expression over the stored text and no model is ever asked to produce a number. Two inputs are model-graded from those same transcripts, a recognition flag and a sentiment value, and together they can move an engine score by at most 30 of 100 points. That split is published on the technology page rather than presented as fully mechanical.

Separating real movement from engine noise

Assume any single change of one or two mentions is noise until the same question set repeats it. These systems are not deterministic, and the pages they retrieve move underneath them, so two identical questions on the same day can return different businesses in a different order. We cover the mechanism in why AI answers change every time you ask.

Four rules that hold up in practice:

  1. Compare rates, never single questions. Named in 3 of 5 against 1 of 5 is a signal. Question four flipping is not.
  2. Require two consecutive runs. A change that appears once and is gone next month was never a change.
  3. Treat a one-engine move carefully. ChatGPT and Gemini disagree constantly. In our study, across the 137 businesses with complete records from both engines, Gemini scored the same business higher than ChatGPT for 105 of them, ChatGPT was higher for 7, and 25 were tied. Across all 137 the two engines' scores differed by an average of 25.3 points out of 100.han ChatGPT 77% of the time, with a mean gap of 25.3 points.
  4. Look for a cause in the world. If nothing changed outside the measurement, be suspicious of a change inside it.

Our own study is a snapshot, not a trend, and we say so on its own page. Each question was asked exactly once, with no repetition runs, so a business one answer short of a mention on the day stays unmentioned in that data. If you are using our figures as a benchmark, use them as a picture of one day rather than as a moving average.

What a monitoring report must contain to be worth paying for

A monitoring report earns its fee when you can check it yourself without asking the vendor a single question. Seven things make that possible.

  1. The exact questions asked, printed in full, not summarised as "a range of buyer queries".
  2. Which engines were asked, and how, including whether web search or search grounding was enabled.
  3. The date of each run.
  4. Every answer stored word for word and readable by you, not just a count.
  5. The result expressed as a rate with its denominator: named in 3 of 5, not "visibility improving".
  6. The scoring formula, published, so the number can be recomputed and argued with.
  7. The limits of the method, written down by the vendor rather than left for you to discover.

We hold ourselves to the last one in a specific way. Our technology page records a live defect in our own reporting: when Gemini's search grounding falls back, the run counts that fallback in its own metadata, but the report page still describes the direct verdict as grounded. On a quota-exhausted run our report therefore reads more confident than the run actually was. It is written on the method page because a limit you have to be told about by someone else was never a disclosed limit.

If you are comparing quotes, what AI visibility costs in Australia sets out the four price bands and what belongs in a fair one.

Get the baseline before you spend anything.

The free scan asks real buyer questions of ChatGPT and Gemini, counts how often you are named, and shows you one real quote of what an engine said about you today. No email gate.

Run my free scan