Skip to main content

Research · Methodology

Perth AI Search Study 2026: methodology

Everything needed to check the report or repeat the study. The question set was frozen and checksummed before any data was collected, and the collection code was committed before the first wave ran.

The question set

200 questions across 20 Perth industries, 10 per industry. Half were five fixed phrasings applied identically to every industry, changing only the service term, so wording could be compared as a variable. Half were natural consumer phrasings written for each industry.

The set was frozen before collection and checksummed. Fifteen automated checks confirmed 200 rows, 20 industries, 10 questions per industry, no duplicate text, Perth intent in every question, each fixed phrasing appearing exactly once per industry, and an identical service term across the five fixed phrasings within each industry.

SHA256 of the frozen question set: 6454e62dade11ff2b36502c4cee3c1c6d0c7d9f02cf259b6bf28fdd4d09eca2b

Platforms and parameters

All data was collected through the DataForSEO Standard API. Models were not selected: on both assistant surfaces the platform decides, and the model it reported is recorded with every observation. Every parameter was identical across all three waves.

Endpoints and parameters
SurfaceEndpointParametersLocation sent
Google AI Overview, plus organic top 10 serp/google/organic/task_post load_async_ai_overview true, se_domain google.com.au, depth 10, desktop Perth, Western Australia
Google AI Mode serp/google/ai_mode/task_post desktop Perth, Western Australia
ChatGPT ai_optimization/chat_gpt/llm_scraper/task_post force_web_search, expand_citations Australia (country level is the finest available)
Gemini ai_optimization/gemini/llm_scraper/task_post expand_citations Perth, Western Australia
Google Maps baseline serp/google/maps/task_post depth 20 Perth, Western Australia

The ChatGPT interface exposes 214 locations, all country level, so Australia is the finest targeting available. Perth intent was carried in the question text instead. This asymmetry against the other three surfaces is a limitation, not a choice, and comparisons are qualified accordingly.

Collection waves

Three waves, separated in time and spanning two parts of the day, so the study measures short-term stability rather than three generations taken seconds apart. Both UTC and Perth time are recorded for every observation.

Collection waves
WavePerth timeUTCSeparation
18 September 2026, 17:062026-09-08 09:06 UTCbaseline
29 September 2026, 13:072026-09-09 05:07 UTC20.0 hours after wave 1
310 September 2026, 16:362026-09-10 08:36 UTC27.5 hours after wave 2

All 3,000 requests returned successfully. There were no failed or partial observations to exclude.

Identifying businesses in an answer

Where a platform returned structured business records, those were used directly. Where a business was named only in prose, it was accepted only if it matched a business in the Perth Maps results for that question, or carried a recognised business-type word. Everything else was discarded and counted rather than guessed at, which under-counts rather than inventing businesses.

A business counts as recommended only when the answer presents it as a suggested option. Appearing only in a citation, in an underlying search result, or in a cautionary sentence counts as a mention, not a recommendation.

Businesses were matched across answers using the Google Business Profile identifier first, then the business's own website domain, then an exact normalised name. A near-name match never merges two businesses automatically: it creates a separate record and flags it for manual review. Shared domains such as social networks and marketplaces are excluded from domain matching, because several unrelated businesses can share one.

Exclusions

Four exclusions were applied after collection, each recorded in the dataset rather than deleted, so any of them can be reversed and checked.

  • Platforms and directories named as though they were businesses. Extraction had accepted Reddit, Facebook, Airtasker and similar as recommended businesses. 31 were excluded.
  • Utilities and regulators referenced for context rather than recommended, such as the state electricity network. 4 were excluded.
  • Six businesses named by the AI systems that could not be corroborated with a Google Business Profile identifier or a confirmed street address anywhere in the collection. They may well be real, but they are not verifiable from this dataset, so they do not appear in any table.
  • Geographic scope was not resolvable for most businesses, so the study makes no claim about local independents versus national brands. Only 7.2% could be confidently classified.

One example of why the third exclusion matters. A plumbing business appearing in Perth Maps results was nearly reported as a Perth local. A national search matched on its own domain returned a second branch in New South Wales, which makes it a national operator. It is recorded as such.

Statistical approach

The study is observational and cross-sectional. No result establishes cause. The unique question is the unit of analysis, and the three repeat runs are treated as repeated measures of the same question rather than as three independent searches.

Because each industry contains one instance of each fixed phrasing with an identical service term, phrasing comparisons are paired within industry. Paired tests were used for them. An unpaired test would treat the same industry's five phrasings as five independent samples, which they are not.

Where businesses were compared, models cluster on the question, because businesses appearing under the same question are not independent of each other. Similarity between answer sets is reported both as a raw overlap measure and as a share of the maximum attainable given the two list lengths, because a system naming three businesses and a system naming twenty cannot reach a high raw overlap even in perfect agreement.

The measured characteristics explain roughly a quarter of the variation between businesses. That figure is published because it is the most important limitation in the study.

Provenance and cost

Published by
AI SEO Perth
Research conducted by
Search Scope
Lead researcher
Dorian Menard
Dataset
Perth AI Search Study 2026
Collection period
8 to 10 September 2026
Methodology version
2.0
Observations
3,000 across five surfaces, 2,400 of them AI answers
Data cost
USD 4.68

Methodology version 2.0 differs from version 1.0 in scope: 20 industries rather than one, 200 questions rather than 25, four platforms rather than five, and collection through an API rather than a browser session.