To get cited in Perplexity, allow PerplexityBot in robots.txt, lead each section with a direct answer, and name your entities explicitly. Perplexity runs a live web search on every query, so retrievability decides eligibility and extractability decides selection.
How Perplexity selects sources
Perplexity is an answer engine, not an index you rank in. Its help documentation describes searching the web in real time for each question, drawing on large language models — currently including GPT-5 and Claude 4.6 Sonnet — to interpret the query and compose a response. Every answer carries numbered citations back to the pages it used.
Three mechanics follow, and they set the agenda for everything below.
There is no training-data fallback. Perplexity retrieves at query time. If its crawler cannot reach your page, you are not eligible, regardless of content quality.
Selection happens after retrieval. Perplexity gathers candidate pages, then cites a subset. Being retrievable is necessary and not sufficient.
Dates are a first-class field. Perplexity's API returns each search result with its URL, title and publication date, so recency is structured data the system handles rather than something it infers from your prose.
What the research actually shows
The most rigorous public test of AI search citation behaviour is the Tow Center for Digital Journalism's March 2025 study, AI Search Has a Citation Problem by Klaudia Jaźwińska and Aisvarya Chandrasekar. The researchers ran 1,600 queries — 20 publishers × 10 articles × 8 chatbots — asking each engine to identify the headline, publisher, date and URL of a supplied excerpt.
| Finding | Figure | Why it matters for your pages |
|---|---|---|
| All eight engines, incorrect answers | More than 60% | Attribution is the weak link in AI search, not retrieval volume. Pages that are easy to attribute have an edge. |
| Perplexity, incorrect answers | 37% | The best of the eight tested - and still wrong on roughly one query in three. |
| Grok 3, incorrect answers | 94% | Engine choice matters more than page-level tactics. Perplexity is the one worth optimising for first. |
| Paid tiers vs free tiers | Higher error rates | Premium models gave definitive wrong answers instead of declining. Ambiguous pages get confidently misattributed. |
Read the 37% the right way. It is not a reason to ignore Perplexity — it is the best-performing engine in the study. It is a reason to make your page trivially attributable: an unambiguous title, a visible publication date, your organisation named in the text rather than only in the logo.
The two crawlers, and which one actually matters
Perplexity operates two agents with different jobs. Confusing them is the most common way a site makes itself uncitable.
| User-agent | Job | If you block it |
|---|---|---|
| PerplexityBot | Builds the search index that makes you eligible to be surfaced and linked | You cannot be cited. This is the one that matters. |
| Perplexity-User | Fetches one page on demand when a user's question requires visiting it | Limited. Perplexity's documentation notes this agent generally ignores robots.txt, because a user initiated the request. |
Perplexity's bot documentation recommends permitting PerplexityBot so your content can be surfaced. Check two layers, not one. A CDN bot-management rule or a one-click "block AI scrapers" toggle rejects the crawler before it ever reads robots.txt, and that is the most common reason a correct robots.txt changes nothing.
Six changes that move the needle
- Allow PerplexityBot, then verify in your logs. Do not assume. Filter your access logs for the user-agent and confirm 200 responses, not 403s. Check the WAF and CDN separately from robots.txt.
- Server-render the substance. Content that appears only after client-side JavaScript runs is content a crawler may never see. The answer must be in the HTML response.
- Lead every section with its answer. Put a direct, self-contained answer in the first sentence under each heading. Perplexity composes from passages; a passage that stands alone is easier to lift than one that needs three paragraphs of setup.
- Name entities explicitly. Write "Perplexity's PerplexityBot crawler" rather than "it". Given the Tow Center's 60% misattribution rate across engines, a passage that is unambiguous out of context is a passage that survives being quoted out of context.
- Put a real date on the page. Perplexity's API carries a publication date per result. Expose accurate published and modified dates in your markup — and never backdate or fake a refresh.
- Use tables for comparisons. Structured rows are extracted more reliably than the same facts buried in prose.
Perplexity vs ChatGPT Search vs Google
| Dimension | Perplexity | ChatGPT Search | |
|---|---|---|---|
| Retrieval | Live search on every query | Live in browsing mode; training data otherwise | Crawled and indexed ahead of the query |
| Citations shown | Numbered, on every answer | Inline, when browsing is used | AI Overview with expandable sources |
| Published ranking criteria | None | None | Extensive public documentation |
| Crawler to allow | PerplexityBot | OAI-SearchBot | Googlebot |
| Citation accuracy (Tow Center, 2025) | 37% incorrect - best of 8 engines tested | Included in the 60%+ overall error rate | Gemini included in the same study |
| Success metric | Citation rate on your questions | Citation rate on your questions | Rank position and click-through |
The third row is the practical point. With no published criteria for two of the three, Perplexity SEO cannot be run off a vendor checklist. It has to be run off measurement.
How to measure whether it worked
Perplexity lists its sources on every answer, which makes the feedback loop tighter than anywhere else in AI search. Treat it as an experiment.
- Write 30 to 100 questions, not keywords. Phrase them the way a buyer would type them. Questions are the input format the engine receives.
- Record the baseline. For each question, note which sources Perplexity cites today and whether you appear. Most brands find they are absent from the majority of their own category's questions.
- Read the pages that are winning. The cited sources are listed for you — competitive intelligence that traditional search never handed anyone this cheaply.
- Change one thing per cycle. Fix crawler access, or restructure one page answer-first. One variable, so the result is attributable.
- Re-measure weekly. Citation sets shift as the index updates, and the Tow Center study found engines give different answers to the same question. A single check can mislead you in either direction; the trend line is the signal.
- Track share, not presence. Appearing in one question out of twenty is a different position from one out of two.
Step 5 is where manual checking breaks down. A hundred questions, repeated weekly, across even one engine is not a task anyone sustains by hand.
Frequently asked questions
Perplexity does not publish its selection criteria. What is documented is the mechanism: a live web search per query, candidate pages retrieved, then an answer composed with numbered citations. Its API exposes each result's URL, title and publication date. Since the weightings are unpublished, make pages retrievable and quotable, then measure your real citation rate.
The Tow Center for Digital Journalism tested eight AI search engines across 1,600 queries in March 2025. Perplexity answered 37 percent incorrectly — the best result of the eight, while the engines collectively erred on more than 60 percent. Paid tiers performed worse, giving confident wrong answers rather than declining.
PerplexityBot. It builds the index that determines whether you are eligible to be surfaced, and Perplexity's documentation recommends permitting it in robots.txt. A second agent, Perplexity-User, fetches pages on demand and generally ignores robots.txt. Check your CDN or WAF too — those block crawlers before robots.txt is ever read.
They share fundamentals — answer-first structure, crawler access, unambiguous entities — but differ in two ways. Perplexity runs a live search on every query with no training-data fallback, so retrievability is non-negotiable. And it cites sources on every answer, making measurement far easier than on engines that collapse their sources.
Three causes cover most cases. Retrieval: PerplexityBot cannot reach the page, usually a CDN or WAF rule rather than robots.txt. Structure: the answer is present but buried where it cannot be lifted as a standalone passage. Specificity: a competitor answers the exact question with concrete facts while your page covers the topic generally.
Yes. Selection happens per question, not per domain, so a focused page answering one question precisely can be cited ahead of a larger site covering the topic broadly. This is the structural advantage smaller sites hold in AI search: you need to be the best source for individual questions, not to outrank anyone across a category.