Skip to main content

Guide

How to measure AI search visibility without rank tracking

Why rank tracking does not transfer to composed answers, what is observable (mentions, recommendations, citations, referrals) and what is not (impressions, query volume, share), how to set a repeated-run baseline with a fixed prompt set, and the measurement mistakes that produce false confidence.

Author
Dorian Menard, founder of Search Scope
Published
Reading time
8 min

Why rank tracking does not transfer

There is no position to track in a composed answer. The answer is generated per session from sources selected at that moment, and running the same prompt an hour later on the same account can return different sources. INFERRED That variance is a property of the systems, not a measurement flaw, and any tool reporting an AI “ranking” as a stable number is relabelling something else.

TESTED The surfaces also move. Pew measured AI summaries on 18% of US Google searches in March 2025; platforms have shipped changes since that altered how sources are surfaced and cited. A baseline therefore has a shelf life: record the date and the model labels it was taken against, or it stops being comparable to the next one.

What is observable

  • Mentions and recommendations. Which businesses an answer names, and which it presents as a shortlist, for a defined prompt run a defined number of times. This is the core observation and it is countable.
  • Citation context. Where a surface shows sources, what was cited and in what role. Being cited as an example of a problem is not the same as being cited as the recommended provider; a count that treats them identically says little. Record what the mention said.
  • Referral traffic. DOCUMENTED Assistants that pass a referrer show up in analytics; GA4 now groups them into an AI Assistant default channel. The numbers are small and the visitors are often well qualified. The GA4 guide covers setup.
  • Google’s AI features, indirectly. DOCUMENTED Google reports AI Overview and AI Mode traffic inside Search Console’s Performance report under the Web search type, not separately. Total web clicks include them; they cannot be isolated.
  • Completed work. Whether access is verified, facts are consistent and markup validates. Binary, checkable, and the most reliable line in any report.

What is not observable

  • Impressions. No platform reports how many people saw a brand named in an answer.
  • Query volume for a category. The denominator is not published for Australia or Perth, so any share figure has been modelled, not measured.
  • Whether a specific model will cite the business tomorrow.
  • Attribution of a sale to an answer. Referrals land on homepages more often than on the page that earned the mention, and most influence is a name read without a click.

INFERRED A stated limit beats a confident guess. A report that says which numbers are solid and which are directional is more useful than one presenting everything at the same confidence.

MeasureObservable?Why
Referral clicks from assistantsYesStandard analytics, AI Assistant channel
Mentions and recommendations per promptYesRepeated-run sampling
Citation contextYes, where shownManual recording per run
Share of category queriesNoDenominator unpublished
Impressions of a brand in answersNoNot logged by any platform
AI Overview clicks in isolationNoFolded into Search Console Web totals

Setting a baseline

  1. Fix the prompt set. Twenty to forty prompts drawn from how buyers actually ask: enquiry emails, call notes, support tickets. “Is it worth doing this for a business with forty staff” is a prompt; “AI SEO services” is a search query someone rewrote. Split across awareness, comparison and decision stages, and include prompts you expect to lose, or the baseline can only move down.
  2. Classify each prompt by query class and buyer stage before running anything, so patterns can be read by class instead of averaged across intents that behave differently.
  3. Control the session. Same platforms, same account conditions, same stated location, every time. Use a clean baseline account; if personalisation is a question, add seeded persona accounts and treat divergence between them as a measurement.
  4. Run each prompt at least three times per platform. A source in two of three runs this round and one of three next round has not necessarily declined.
  5. Record everything per run: date, time, platform, model label, brands mentioned, brands recommended and their order, citations, source classes, notes. The publication’s observation schema is a ready-made record structure.
  6. Freeze the set. Changing prompts breaks comparability. Add a new set alongside if the business changes; do not edit the old one.

Read the baseline for direction across rounds, not for movement within one.

Common measurement mistakes

  • Reporting a share. “We appear in a third of category prompts” implies a known denominator. Report “in 11 of 30 prompts, across 3 runs each, on this date, on this platform” instead.
  • One run. A single answer is an anecdote in either direction.
  • Ignoring location. An answer generated in Perth can differ from one generated through an east-coast connection. Record the region and the location setting for every run.
  • Ignoring context. A negative mention counted as a positive one.
  • Chasing one prompt. Individual prompts drop in and out for reasons unrelated to the site. Structural causes are worth fixing; individual appearances are not worth chasing.
  • Merging rounds across model changes. Note the model label; a change in it is a break in the series.

What a report should contain

In order: work completed, with evidence; prompt-set movement, with run counts shown so the reader can judge the sample; referral and assisted-conversion data, with the under-count stated; and a plain list of what remains unknown. Prompt movement is a quarterly item; monthly reporting of it produces noise.

The honest limit

Measurement in this field is closer to survey research than to analytics: samples, distributions, wide uncertainty. Directional evidence is still evidence, and enough to make budget decisions with, provided the size of the investment matches the uncertainty and the report says what that uncertainty is.

What is documented and what is inferred

  • Documented: Search Console’s treatment of AI feature traffic; GA4’s AI Assistant channel; Pew’s frequency figure.
  • Inferred: the reasons for variance; the baseline design; the reporting order.
  • Not established: Perth query volumes, and any Perth frequency until the observation programme publishes.

Sources

  1. AI features and your website, Google Search Central (read 9 September 2026) DOCUMENTED

    AI feature traffic is reported inside the Search Console Performance report, Web search type.

  2. Default channel group, Google Analytics Help (read 9 September 2026) DOCUMENTED

    Defines the AI Assistant channel (medium 'ai-assistant').

  3. Google users are less likely to click on links when an AI summary appears in the results, Pew Research Center, 22 July 2025 (read 9 September 2026) TESTED