Research · Methodology
Perth AI Search Study 2026: methodology
Everything needed to check the report or repeat the study. The question set was frozen and checksummed before any data was collected, and the collection code was committed before the first wave ran.
The question set
200 questions across 20 Perth industries, 10 per industry. Half were five fixed phrasings applied identically to every industry, changing only the service term, so wording could be compared as a variable. Half were natural consumer phrasings written for each industry.
The set was frozen before collection and checksummed. Fifteen automated checks confirmed 200 rows, 20 industries, 10 questions per industry, no duplicate text, Perth intent in every question, each fixed phrasing appearing exactly once per industry, and an identical service term across the five fixed phrasings within each industry.
SHA256 of the frozen question set:
6454e62dade11ff2b36502c4cee3c1c6d0c7d9f02cf259b6bf28fdd4d09eca2b
Platforms and parameters
All data was collected through the DataForSEO Standard API. Models were not selected: on both assistant surfaces the platform decides, and the model it reported is recorded with every observation. Every parameter was identical across all three waves.
| Surface | Endpoint | Parameters | Location sent |
|---|---|---|---|
| Google AI Overview, plus organic top 10 | serp/google/organic/task_post | load_async_ai_overview true, se_domain google.com.au, depth 10, desktop | Perth, Western Australia |
| Google AI Mode | serp/google/ai_mode/task_post | desktop | Perth, Western Australia |
| ChatGPT | ai_optimization/chat_gpt/llm_scraper/task_post | force_web_search, expand_citations | Australia (country level is the finest available) |
| Gemini | ai_optimization/gemini/llm_scraper/task_post | expand_citations | Perth, Western Australia |
| Google Maps baseline | serp/google/maps/task_post | depth 20 | Perth, Western Australia |
The ChatGPT interface exposes 214 locations, all country level, so Australia is the finest targeting available. Perth intent was carried in the question text instead. This asymmetry against the other three surfaces is a limitation, not a choice, and comparisons are qualified accordingly.
Collection waves
Three waves, separated in time and spanning two parts of the day, so the study measures short-term stability rather than three generations taken seconds apart. Both UTC and Perth time are recorded for every observation.
| Wave | Perth time | UTC | Separation |
|---|---|---|---|
| 1 | 8 September 2026, 17:06 | 2026-09-08 09:06 UTC | baseline |
| 2 | 9 September 2026, 13:07 | 2026-09-09 05:07 UTC | 20.0 hours after wave 1 |
| 3 | 10 September 2026, 16:36 | 2026-09-10 08:36 UTC | 27.5 hours after wave 2 |
All 3,000 requests returned successfully. There were no failed or partial observations to exclude.
Identifying businesses in an answer
Where a platform returned structured business records, those were used directly. Where a business was named only in prose, it was accepted only if it matched a business in the Perth Maps results for that question, or carried a recognised business-type word. Everything else was discarded and counted rather than guessed at, which under-counts rather than inventing businesses.
A business counts as recommended only when the answer presents it as a suggested option. Appearing only in a citation, in an underlying search result, or in a cautionary sentence counts as a mention, not a recommendation.
Businesses were matched across answers using the Google Business Profile identifier first, then the business's own website domain, then an exact normalised name. A near-name match never merges two businesses automatically: it creates a separate record and flags it for manual review. Shared domains such as social networks and marketplaces are excluded from domain matching, because several unrelated businesses can share one.
Exclusions
Four exclusions were applied after collection, each recorded in the dataset rather than deleted, so any of them can be reversed and checked.
- Platforms and directories named as though they were businesses. Extraction had accepted Reddit, Facebook, Airtasker and similar as recommended businesses. 31 were excluded.
- Utilities and regulators referenced for context rather than recommended, such as the state electricity network. 4 were excluded.
- Six businesses named by the AI systems that could not be corroborated with a Google Business Profile identifier or a confirmed street address anywhere in the collection. They may well be real, but they are not verifiable from this dataset, so they do not appear in any table.
- Geographic scope was not resolvable for most businesses, so the study makes no claim about local independents versus national brands. Only 7.2% could be confidently classified.
One example of why the third exclusion matters. A plumbing business appearing in Perth Maps results was nearly reported as a Perth local. A national search matched on its own domain returned a second branch in New South Wales, which makes it a national operator. It is recorded as such.
Statistical approach
The study is observational and cross-sectional. No result establishes cause. The unique question is the unit of analysis, and the three repeat runs are treated as repeated measures of the same question rather than as three independent searches.
Because each industry contains one instance of each fixed phrasing with an identical service term, phrasing comparisons are paired within industry. Paired tests were used for them. An unpaired test would treat the same industry's five phrasings as five independent samples, which they are not.
Where businesses were compared, models cluster on the question, because businesses appearing under the same question are not independent of each other. Similarity between answer sets is reported both as a raw overlap measure and as a share of the maximum attainable given the two list lengths, because a system naming three businesses and a system naming twenty cannot reach a high raw overlap even in perfect agreement.
The measured characteristics explain roughly a quarter of the variation between businesses. That figure is published because it is the most important limitation in the study.
Provenance and cost
- Published by
- AI SEO Perth
- Research conducted by
- Search Scope
- Lead researcher
- Dorian Menard
- Dataset
- Perth AI Search Study 2026
- Collection period
- 8 to 10 September 2026
- Methodology version
- 2.0
- Observations
- 3,000 across five surfaces, 2,400 of them AI answers
- Data cost
- USD 4.68
Methodology version 2.0 differs from version 1.0 in scope: 20 industries rather than one, 200 questions rather than 25, four platforms rather than five, and collection through an API rather than a browser session.