What we did
We asked four AI systems 200 commercial questions covering 20 Perth industries, from “best plumber Perth” to “best buyers agent in Perth for first home buyers”. We asked every question three times, on three separate days, and recorded every answer. That is 2,400 AI answers, collected between 8 and 10 September 2026, at a total data cost of USD $4.68.
For each answer we recorded which businesses were named, which were actually recommended rather than merely mentioned, and which web pages the system cited. We recorded Google’s ordinary organic results and Google Maps results for the same questions, so we could compare. The Google Maps baseline is a fifth surface, captured once per question per round, which is another 600 records. The full collection is 3,000 observation records: 2,400 AI answers across four systems, and 600 Maps baseline captures.
This report is published by AI SEO Perth, a research publication operated by Search Scope. Search Scope is an SEO agency in Perth, and “SEO agency” is one of the 20 industries we measured. Search Scope does not appear in the top 40 most-recommended businesses on any of the four platforms. We are stating that at the top rather than the bottom.
This page exists as evidence, not because anyone searches for it. We did not choose this topic from keyword demand and we are not claiming it has any.
Most recommended businesses were not the Maps winners
Between 60% and 78% of the businesses these systems recommended did not appear in the Google Maps top 20 for the same question.
That was the first result that surprised us. The assumption going in was that an AI assistant asked for a local plumber would mostly reheat the local pack. It did not.
| System | Recommendations not in the Maps top 20 |
|---|---|
| Google AI Overview | 77.6% |
| Gemini | 72.1% |
| ChatGPT | 64.7% |
| Google AI Mode | 60.1% |
Maps position still mattered. Businesses ranking 1 to 3 in Maps were recommended more often than those ranking 11 to 20 on every platform. It simply did not decide the outcome, and a large majority of what these systems named came from outside that list.
The answers kept changing
We asked each question three times over three days. Most recommendations did not survive.
| System | Businesses recommended in all three rounds |
|---|---|
| Google AI Mode | 31% |
| ChatGPT | 17% |
| Google AI Overview | 11% |
| Gemini | 8% |
This matters before any discussion of mechanism. A business owner checking whether ChatGPT recommends them is taking a single reading of a number that moves. One check is not a measurement.
Some of the AI Overview instability is a different thing again. An AI Overview did not always appear. Of the 200 questions, 91 produced one in all three rounds, 71 produced one in some rounds and not others, and 38 never produced one at all. So roughly a third of questions were unstable at the level of whether an AI answer appeared, before anyone looks at what was inside it.
No business owned a platform
Instability in single answers does not say whether one or two businesses were quietly winning most of the time, so we measured concentration as well. For each business and each system we counted the questions on which it was recommended in at least one of the three rounds, and divided by the questions on which that system produced usable recommendation data. We also counted it per run, recommended answers over usable answers, because that is how Search Scope defines its Recommendation Share metric in client reporting.
| System | Questions with usable data | Highest share of questions | Highest share of runs |
|---|---|---|---|
| Google AI Mode | 197 | 10.2% | 8.9% |
| ChatGPT | 200 | 5.5% | 4.9% |
| Google AI Overview | 152 | 3.9% | 3.0% |
| Gemini | 184 | 3.8% | 4.1% |
The ceiling was Bodysmart Physio Pilates & Chiro Perth on Google AI Mode, recommended on 20 of 197 questions. On the other three systems no business reached 6% by either count. Platforms, directories and utilities are excluded from this table, along with the six candidates our manual check could not tie to a Google business identifier and a Perth street address. The exclusions are flags in the dataset, not deletions.
Underneath the noise, two things persisted
This is the part we think matters most, and it only becomes visible once you accept the instability above rather than averaging it away.
We sorted every business into five levels, from never recommended, through recommended once, up to recommended repeatedly across three or four different AI systems. Then we looked at what changed as businesses moved up.
| Never | Once | Repeated, one system | Two systems | Three or four systems | |
|---|---|---|---|---|---|
| Businesses | 2,542 | 586 | 425 | 770 | 217 |
| In Google organic top 10 | 1.5% | 8.2% | 9.4% | 26.9% | 52.1% |
| Median Google reviews | 21 | 62 | 78 | 135 | 234 |
| Median star rating | 4.90 | 4.90 | 4.90 | 4.90 | 4.90 |
| Median Maps position | 12 | 9 | 9 | 7 | 5 |
Two columns climb steadily. Organic presence goes from 1.5% to 52.1%. Median reviews go from 21 to 234. Star rating does not move at all.
Holding Maps position and star rating constant and comparing businesses within the same question, the half with more reviews was recommended 68.5% of the time against 37.3% for the half with fewer. That held inside every Maps ranking band.
Getting in is not the same as staying in
Presence on the third-party pages these systems cite, the directories and review sites, was associated with being recommended once. It was not associated with being recommended repeatedly.
| Stage | Third-party page presence | Organic top 10 |
|---|---|---|
| Never recommended, to recommended at all | 1.88 | 6.93 |
| Recommended once, to recommended repeatedly | 1.35 | 2.89 |
| One system, to more than one | 1.33, not significant | 3.54 |
| Two systems, to three or four | 0.76, not significant | 3.00 |
Those figures are odds ratios. Above 1 means more likely, below 1 means less likely.
The pattern worth noticing is that directory presence fades as you climb and organic presence does not. We tried to break this result before accepting it. The obvious explanation would be that the sample thins out at the top and the statistics simply lose power. It is not that: at the final stage, third-party presence had more businesses behind it than organic presence did, and it still showed no association while organic held.
Maps position and star rating also stopped being associated at that final stage. Only organic presence and review count carried through.
What did not work
Three things we expected to matter, and did not.
A perfect five-star rating. Businesses rated 4.7 to 4.9 were recommended 59.5% of the time. Businesses rated 4.9 to 5.0 were recommended 49.6% of the time. The likely reason is not a penalty on perfection: five-star businesses in our sample had roughly half as many reviews, a median of 38 against 89. Volume appears to be doing the work, and a perfect score usually signals a small number of reviews.
Position on a directory listing. Among businesses appearing on the cited pages we fetched, those listed in positions 4 to 10 were recommended slightly more often than those in positions 1 to 3. There was no clean relationship in either direction.
“Best of” and “top rated” badges. Businesses carrying one on a cited page were recommended 28.9% of the time. Businesses without one, 27.9%. Effectively no difference.
The four systems are not one thing
Google’s two AI products disagreed with each other more than we expected. Seven in ten businesses named by an AI Overview appeared nowhere in Google’s AI Mode answer to the identical question, even though AI Mode typically named around twenty businesses and had ample room to include them.
They also appear to draw on different material. Of the businesses each recommended, restricted to those we could match to a website:
| In Google organic top 10 | In Maps top 20 | |
|---|---|---|
| Google AI Overview | 66.8% | 37.5% |
| Google AI Mode | 22.7% | 78.6% |
That pattern held in all 20 industries. It survived our attempts to explain it away as an artefact of list length, of AI Overviews appearing less often, or of directories sitting inside organic results.
The citation behaviour splits the same way. ChatGPT cited Google zero times across 4,456 citations. Google’s AI Mode cited little else, at 85.8%, almost entirely Google-hosted business panels. AI Overviews leaned on businesses’ own websites at 38.2%. Gemini leaned on business websites too, at 42.0%.
Asking a question worked better than typing keywords
For 100 of our 200 questions we used five fixed phrasings per industry, changing only the wording and keeping the service and the location identical. The proportion that produced an AI Overview:
| Phrasing | Produced an AI Overview |
|---|---|
| “best [service] Perth” | 30% |
| “top [service] Perth” | 50% |
| “recommended [service] Perth” | 60% |
| “who is the best [service] in Perth” | 95% |
| “recommend a good [service] in Perth” | 100% |
Keyword-style phrasings produced an AI Overview for 46.7% of questions. Question-style phrasings, 97.5%. Across every industry, there was not one case where a keyword phrasing produced an AI Overview and its matching question phrasing did not.
One honest caveat: the question phrasings are also longer, and we cannot fully separate length from grammar. We can say length alone does not explain it, because “best [service] Perth” and “top [service] Perth” are the same length and 20 percentage points apart.
What we cannot tell you
The characteristics we measured explain roughly a quarter of the variation between businesses. Put plainly, around three-quarters of what decides whether a particular AI system recommends a particular business on a particular day was not captured by review count, star rating, Maps position, organic ranking or directory presence.
We did not measure website structure, schema markup, llms.txt files, page speed, Google Business Profile completeness, brand mentions or anything about the content of business websites. We have nothing to say about them, in either direction.
Nothing here is causal. Every result is an association observed at one point in time. A business with more reviews was more likely to be recommended in this sample. That is not the same as saying more reviews would cause an AI system to recommend it.
Provenance
Published by: AI SEO Perth
Research conducted: Search Scope
Lead researcher: Dorian Menard
Dataset: Perth AI Search Study 2026
Collection period: 8 to 10 September 2026
Methodology version: 2.0
The full method, including the frozen question set and its checksum, every API endpoint and parameter, the entity matching rules, the exclusions and the statistical tests, is on the methodology page. The complete observation dataset is published alongside it.