Skip to main content

Reference

Perplexity SEO: how Perplexity retrieves and cites sources

Author
Dorian Menard, founder of Search Scope
Published
Reading time
5 min

Perplexity SEO is the work of being reachable by, and cited in, Perplexity's answers. Perplexity documents two agents: PerplexityBot, which surfaces and links websites in its search results, and Perplexity-User, which visits a page when a user asks a question and includes a link to it in the response. Perplexity states neither collects content for training foundation models. Because Perplexity shows citations by default, it is the surface where source patterns can be observed most completely.

Key facts

Key facts about Perplexity SEO
FactDetailEvidence
Search index agentPerplexityBot is designed to surface and link websites in search results on Perplexity and is not used to crawl content for AI foundation models. DOCUMENTED
User-triggered fetchPerplexity-User might visit a web page to help provide an accurate answer and include a link to the page in its response. It generally ignores robots.txt because a user requested the fetch. DOCUMENTED
TrainingPerplexity states neither agent is used to collect content for training AI foundation models. DOCUMENTED
FirewallsPerplexity publishes IP ranges and says sites using a web application firewall may need to explicitly allow its bots. DOCUMENTED
Source selectionNot documented. How Perplexity ranks retrieved pages and picks citations is not stated in its crawler documentation. DOCUMENTED

What Perplexity documents

DOCUMENTED Perplexity’s crawler documentation names two agents and is explicit about what each does.

AgentPerplexity’s stated purposeNotes from the documentation
PerplexityBot“Designed to surface and link websites in search results on Perplexity.”“Not used to crawl content for AI foundation models.” Perplexity recommends allowing it in robots.txt and allowing its published IP ranges.
Perplexity-User“When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a link to the page in its response.”“Not used for web crawling or to collect content for training AI foundation models.” “Since a user requested the fetch, this fetcher generally ignores robots.txt rules.”

DOCUMENTED Perplexity adds a practical note most platforms leave out: “If you’re using a Web Application Firewall (WAF) to protect your site, you may need to explicitly whitelist Perplexity’s bots”, with configuration guidance for common providers. For a Perth business behind a managed firewall or a CDN with bot rules, this is the most likely reason to be absent from Perplexity answers, and it is fixable in a configuration screen rather than in content.

Why Perplexity matters for research more than for volume

DOCUMENTED Perplexity’s documentation ties the user-triggered fetch to a visible link in the response. In practice the product shows numbered citations on its answers by default, which makes it the surface where citation patterns can be recorded most completely: every source used to compose an answer is more likely to be visible than on surfaces that show a subset of links.

INFERRED That is why this publication’s methodology treats Perplexity as a diagnostic surface. A Perth prompt run on Perplexity three times yields a near-complete list of the pages that were composed from, classified by source type. The same prompt on a surface that shows fewer citations yields a partial list. The Perplexity result is not a proxy for the others, but it is the clearest picture of what a retrieval engine finds for a Perth question.

No claim is made here about how many people in Perth use Perplexity, because no source read for this page measures it.

What is not documented

Perplexity does not publish how it ranks the pages PerplexityBot has indexed, how many pages it reads per question, or how it chooses which to cite among those read. This publication does not assert a ranking formula or an index provider for Perplexity. Anything more specific than the table above is either an observation with a date and a run count, or a hypothesis.

What a Perth business can check today

  1. Access. Confirm that PerplexityBot is allowed in robots.txt and that the firewall or CDN does not block Perplexity’s published IP ranges. The crawler access guide covers checking the raw HTML the agent receives.
  2. Retrieval. Ask a Perth question the business should answer, three times in fresh threads, and record the cited pages. If the business’s page is absent while a directory or competitor page is cited, the useful question is what those cited pages state that the business’s page does not.
  3. Facts elsewhere. Because Perplexity composes from several sources, stale facts on third-party pages surface in its answers. The entity consistency audit is the fix.

How Perplexity applies in Perth

HYPOTHESIS The planned Perplexity Perth source study records every citation for a fixed Perth prompt set, reduces it to domains, classifies each domain as owned, third-party, platform-owned, UGC, directory, publisher or brand source, and reports the share of citations in each class. The hypothesis it tests is that directories, review platforms and national publishers make up most citations for Perth recommendation prompts. That is stated as a hypothesis, not a finding.

What is documented and what is inferred

  • Documented: the two agents and their purposes; the training statement; robots.txt behaviour of Perplexity-User; firewall guidance.
  • Inferred: Perplexity’s value as a diagnostic surface because of citation visibility.
  • Hypothesis: the Perth source-class mix.
  • Undisclosed: ranking and citation selection among retrieved pages; usage figures for Perth.

Sources

  1. Perplexity crawlers, Perplexity (read 9 September 2026) DOCUMENTED