Skip to main content

Data

Data and downloads

Everything the programme publishes as a file. Datasets appear here only once observations exist. The schema, the methodology and the project registry are published now so the structure of future data is fixed in public before any of it is collected.

Published datasets

Structure and methodology files

Schema

Observation record schema

JSON Schema for one observation: prompt, class, platform, run, session profile, brands mentioned and recommended, citations, source classes, capture and methodology version.

Updated 2026-09-09

Methodology

Methodology v1.0

The observation methodology as data: query classes, platforms, session profiles, recorded fields, source classes, evidence labels, limitations and version history.

Updated 2026-09-09

Methodology

Research project registry

Every research project with its status, platforms, query classes and dataset links. Planned projects carry no dataset.

Updated 2026-09-09

Example only

Perth AI search observations: example file

Two placeholder rows showing the record structure in both download formats. Every value is a stand-in.

Placeholder values that illustrate the structure. Not observations. Do not cite.

What one observation record contains

The example rows below use placeholder values so the structure can be seen without any real observation being implied.

Example rows. Placeholder values shown to illustrate the record structure. Not observations.

Example observation rows with placeholder values.
ID Date Platform Class Prompt Recommended Cited domains Run
example-project-chatgpt-p001-20260101-r01 2026-01-01 ChatGPT recommendation Example prompt: who are the best example-service providers in Perth? Example Business A, Example Business B example.com.au, directory.example 1
example-project-perplexity-p001-20260101-r01 2026-01-01 Perplexity recommendation Example prompt: who are the best example-service providers in Perth? none named no citations 1
Example observation rows with placeholder values.
Observation record fields
FieldTypeDescription
observation_id string Stable identifier: <project>-<platform>-<prompt-id>-<YYYYMMDD>-r<run>. Example: chatgpt-perth-chatgpt-p001-20261001-r01.
research_project string Slug of the research project the observation belongs to (matches the research collection entry).
prompt string The exact prompt text submitted, unedited.
prompt_id string Identifier of the prompt within the project's fixed prompt set, so repeated runs of the same prompt can be grouped.
query_class enum informational, commercial, comparison, recommendation, local, brand, competitor, trust-validation
buyer_stage enum awareness, consideration, decision, retention, not-applicable
industry string Industry slug from the observatory list (dentists, lawyers, mortgage-brokers, accountants, plumbers, electricians, builders, solar, real-estate, seo-agencies) or 'general'.
location string Geographic context named in the prompt or the session, e.g. 'Perth, Western Australia'.
platform enum chatgpt, google-ai-overviews, google-ai-mode, gemini, perplexity, claude, copilot
model_or_surface string The model or surface as the platform labelled it at the time, e.g. 'GPT-5 (web, search on)', 'AI Overview on google.com.au'.
observation_date string
observation_time string Local time of the run, HH:MM.
timezone string
run_number integer Sequence of this repeat within the observation round. Repeated runs are how session variance is measured.
session_profile enum Which controlled session the run used. Divergence between profiles on the same prompt is itself a measurement.
brands_mentioned array Every business or brand named anywhere in the response.
brands_recommended array Brands the response presented as a recommendation or shortlist entry, a subset of brands_mentioned.
brand_order array brands_recommended in the order the response listed them.
citations_present boolean
cited_urls array
cited_domains array Registrable domains of cited_urls, deduplicated, lower case.
source_classes array Classification of each cited domain.
response_notes string Free-text notes on the response: format, hedging, refusals, non-Perth results.
region_or_session_notes string Anything about location handling, personalisation or logged-in state that could affect the run.
source_capture string Path or URL of the screenshot or raw capture for this run.
methodology_version string Methodology version the run was collected under, e.g. '1.0'.
evidence_type string

Reuse

Published datasets may be quoted and reused with attribution to AI SEO Perth and a link to the report they come from. Observations are dated and versioned; quote the observation dates and methodology version with any figure. Example files must not be quoted as data.