How Perplexity picks its sources

Perplexity retrieves first and writes second. That single fact explains most of what makes a page citable, and most of what keeps one out.

Perplexity picks its sources by running a live search for your question, retrieving a short list of pages, and then writing an answer in which every claim points back to a numbered entry in that list. The practical consequence is simple: a page that is not retrieved cannot be cited, and a page that is retrieved but hard to quote will be listed without being used. Both failures look identical from your side, and both are fixable.

A source list is the numbered set of pages an AI engine retrieved before writing its answer. Perplexity shows it under every response. Appearing in it is the closest thing generative search has to a ranking.

What happens between the question and the answer?

Three steps, in this order.

  1. Retrieval. The engine turns the question into one or more searches and pulls back a working set of pages.
  2. Reading. It extracts passages from those pages that look like they answer the question.
  3. Writing. It composes an answer from the extracted passages and attaches a citation to each claim.

Classic SEO only ever optimised for step one. Steps two and three are where most sites lose, because a page can rank perfectly well and still be unquotable: the answer is buried under an introduction, spread across five paragraphs, or phrased so vaguely that no single sentence can be lifted out of it.

Why is the source list more important than the position?

Because there is no page two. A classic results page has ten links and a scroll bar; a generative answer has a handful of sources, and the reader typically reads the prose rather than the list. Being source number seven of seven is closer to being invisible than to being seventh in Google.

This is why measuring generative visibility as a yes or no per question, repeated over time, is more useful than chasing a rank number. Either you were in the set for that question or you were not.

What makes a page easy to quote?

The pages that get quoted tend to share the same handful of properties, and none of them are exotic:

  • A direct answer in the opening. One or two sentences that stand on their own, near the top, in the form "X is Y" or "Yes, because Z". If a passage needs the paragraph above it to make sense, it is hard to quote.
  • Headings that are the question. A heading phrased as the query gives the engine an obvious anchor for matching a section to a question.
  • Facts with edges. Numbers, dates, versions, named sources, defined terms. A sentence containing a specific figure survives extraction; a sentence containing "significantly improves" does not.
  • Structure the parser can see. Real lists, real tables, real headings, in a clean H1 to H6 hierarchy, rather than styled divs that look like structure to a human and like soup to a machine.
  • Attribution. Links out to primary sources, an identified author, a visible last-updated date. These are the signals an engine has available when it has to choose between two pages that say the same thing.
  • Reachability. Content present in the server-rendered HTML rather than assembled by JavaScript after load, and a robots.txt that does not block PerplexityBot.

That last one is worth stating plainly, because it is the only item on the list that can undo all the others by itself.

How does myGEOscore measure this?

The Perplexity check in the paid report does the same thing you would do by hand, with the guesswork removed.

It reads the page and derives a query set from the strongest signals it can find: the title, the H1, the meta description and the opening text. From those it builds five queries in four shapes, two informational without your brand, one navigational with it, one comparison and one long tail. Each query goes to Perplexity's sonar model, and the citation list that comes back is matched against your site.

A citation counts as yours when the URL matches your full URL, your domain or your hostname, so a link to www.example.com or to a deeper page on the same domain both register. The position of the first match is kept as well, because the difference between first source and last source is real.

The same pattern runs against two other engines: the ChatGPT check reads the url_citation annotations returned by OpenAI's web search tool, and the Google AI Overviews check reads the grounding sources behind a Gemini answer. Together they are the AI Platform Visibility category, and they are reported on their own, beside the score rather than inside it. No page-type rubric gives them any weight, so they move the grade in neither direction. The score measures whether your page is citable; the checks measure whether it was cited. They run in the paid report only. How each check is built, query by query, is written up on the methodology page.

Which factors actually move it?

Visibility is the outcome. The causes are the 48 factors that feed the score, and how much each one counts depends on what kind of page you gave us. A homepage is not weighted like an article. On an article, the weighting is:

  • Content structure and extractability, 25%
  • Authority and trust, 25%
  • Fact density and citability, 20%
  • Structured data and schema, 15%
  • Technical AI readiness, 10%
  • Brand entity and mentions, 5%

The methodology page prints the same table for all five page types. Read the article column next to the properties above and the overlap is the point: content structure is first because extraction happens before ranking matters. Each factor has its own guide explaining what is checked and how the score is produced, in the factor library.

How do you test your own pages?

Do the manual version before you buy anything.

  1. Pick the page you most want cited, and write five questions it should be the answer to. Keep your brand out of four of them.
  2. Ask each one, then read the source list rather than the prose.
  3. Note whether your domain appears and where. Then note which page of yours appeared, since it is often not the one you expected.
  4. Read the passage the engine quoted, if it quoted you. That sentence is the one your page is being judged on.
  5. Repeat in a month. Retrieval sets shift as pages are re-crawled, so one run is a snapshot and three runs are a trend.

The free check covers the 36 factors that decide whether a page is quotable at all, with no signup, and it takes about a minute. The paid report at $19 runs the live citation checks against Perplexity, ChatGPT and Google AI Overviews, and repeats the whole analysis for 3 monthly re-scans so drift shows up as a change rather than a surprise. The pricing page has the full comparison.

Frequently asked questions

Does Perplexity use Google's index?

Perplexity runs its own live retrieval and publishes the source list under each answer. What matters for your page is not which index it came from but whether the page was reachable, retrievable and quotable at the moment of the question.

Why does my page get listed as a source but never quoted?

Usually because nothing on it is a clean, self-contained answer. Retrieval got you into the set; extraction is a separate step. Put a direct answer in the first two sentences under a heading phrased as the question and check again.

Does blocking PerplexityBot hurt me?

It removes you from consideration entirely. A blocked crawler is a binary failure: no amount of content quality compensates. Check robots.txt and your CDN separately, since a CDN can override a robots.txt you never changed.

How often should I re-check?

Monthly is a reasonable cadence for most sites. Retrieval sets change as pages are re-crawled and as competitors publish, and a monthly cadence is short enough to catch a drop while it is still explainable.

Is a citation worth more than a click?

They are different things. A citation puts your name in the answer the reader reads; a click puts the reader on your page. Generative search produces far more of the first than the second, which is why it needs to be measured on its own terms.

Check Your GEO Score

Run a free analysis on your website and see how you score across all 51 factors.

Analyze My Site