Answer
Does schema markup help with AI search?
The short answer
Schema markup helps machines read your facts and is worth doing, but it is not the lever most people sell it as. In our scan of 300 Australian businesses the 229 sites with JSON-LD averaged 41.4 out of 100 against 28.4 for the 71 without. Yet site readiness barely tracked per-engine visibility: Pearson r = 0.116 across the 137 with complete two-engine records.
What our own 300 sites show, with the denominator
Sites with JSON-LD scored higher, and the gap is not small. Of the 300 Australian businesses we scanned on 30 July 2026, 229 had at least one JSON-LD block on the homepage and 71 did not. That is the 24% with no structured data published in the study.
| Group | Sites | Mean AI visibility |
|---|---|---|
| Has JSON-LD on the homepage | 229 | 41.4 |
| No JSON-LD | 71 | 28.4 |
Recomputed from data.csv. Counted over all 300 businesses scanned.
Narrow to the 137 businesses with complete records from both engines and the same direction holds. Of the 98 with schema, 36 were never named by either engine, which is 37%. Of the 39 without, 20 were never named, which is 51%.
One honest note about that first table. Our AI visibility score folds site readiness into itself, and structured data is 25 of the 100 readiness points, so a site with JSON-LD gets a small mechanical head start in that column. The never-named comparison on the 137 does not have that problem, because it counts only whether an engine said the name.
The counterweight nobody publishes: readiness barely predicts naming
A tidy site is not a recommended business. Across the 137 with complete records from both engines, the correlation between site readiness and the mean of the two per-engine visibility scores is a Pearson r of 0.116.
The average site in the sample scored 74.9 out of 100 for readiness. The average business behind those sites scored 38.3 out of 100 for AI visibility. Thirty four of the 137 scored a perfect 100 for readiness, and nine of them were never named by either engine.
We publish that figure knowing exactly what it costs us. It is the number that says the cheapest, most sellable part of this work is not the part that decides the outcome.
The four schema types actually worth adding, and in what order
Start with who you are, then where you operate, then what you sell, then the questions you answer. Anything beyond those four is refinement.
- Organization, with sameAs. Your name, logo, URL, and a sameAs list pointing at your Google Business Profile, directory listings and social profiles. This is the block that lets a machine connect the scattered mentions of you into one identity.
- LocalBusiness, if you serve a place. Address, phone, opening hours, areaServed. Every fact a buyer question turns on, stated once in a form that cannot be misread.
- Service or Product. One entry per thing you actually sell, named the way a customer would name it rather than the way your industry does.
- QAPage or FAQPage, on pages that answer one real question. Only where the page genuinely answers it. Marking up a page that does not answer the question is how a site ends up looking like spam to both machines and people.
Two rules that matter more than the list. The markup must agree with the visible page, and the same facts must agree with what every other site says about you. Structured data that contradicts your own text is worse than none.
Why schema is necessary and not sufficient
Because it changes how easily a machine can read your facts, and not how much the rest of the web corroborates them.
Think of the answer as a claim an engine has to be willing to make in public. Structured data makes your version of the facts unambiguous, which removes a reason to skip you. It does nothing about the reason engines most often name someone else, which is that someone else is written about in more places. In our own sample, 25 of the 137 with complete records were large anchor brands and only two of those were never named, against 54 of the 112 non-anchors.
So the honest sequence is: make the facts legible, then earn the corroboration. Doing the first without the second is a well-marked-up site nobody cites.
The 40% citation boost claim, and why we will not repeat it
We cannot trace it to a study, so we will not use it. You will see a figure of roughly 40% more citations from adding schema quoted across agency blogs and sales decks, usually with no sample size, no method and no link to anything that could be checked.
When a number is repeated without a denominator, the right response is to ask three questions: how many sites, measured how, over what period. If those cannot be answered, the number is marketing. That test disqualifies the 40% claim, and it is the same test we invite you to apply to every figure on this page, which is why each of ours carries its denominator and links to the raw data.
The defensible version of the same idea is much duller. Structured data is associated with higher scores in our sample, the effect of readiness overall on being named is weak, and it costs a few hours. Do it because it is cheap and correct, not because someone promised you 40%.
Where to start
If you want to know whether your own markup is even being read, the free scan checks the homepage for JSON-LD along with robots.txt, the meta basics and llms.txt, then shows you what ChatGPT and Gemini actually say.
Keep reading
Related questions
Answer
Do I need an llms.txt file?
No. Google has said its search systems do not use it.
Read the answer →Answer
Should I block AI crawlers?
What blocking cost the 39 sites in our scan that did it.
Read the answer →Research
AI visibility statistics
The 2026 Australian numbers, each with its denominator.
Read the answer →Learn
The Receipts Engine
How readiness is scored, published point by point.
Read the answer →See what AI says about your business.
Start with a free scan. It is the fastest way to find out whether AI recommends you or your competitor, and you keep the report.