·
Skip to main content

The GEO journal9 min read

How AI engines choose their sources

ChatGPT, Perplexity, Gemini and Claude each document their own crawlers and their own citation behaviour, so ranking first on Google is neither sufficient nor required for being cited. This article separates what the vendors document, what recently published citation studies measured, and what we infer from both — every claim verified on 4 September 2026.

What the vendors actually document

Each of the four engines publishes documentation about its crawlers, and each documents different behaviour. That is why the same question, asked on two engines the same day, returns different cited sources: the retrieval systems differ, by design. Throughout this article we keep three registers apart — what the vendor documents, what we observe in the reports we generate, and what we infer — and we label each as we go.

Here is what the documentation states, verified on 4 September 2026. The table is the documented register only: where a cell reads not documented, we found no public statement from the vendor, and we say so rather than fill the gap.

EngineDocumented crawlersDocumented search behaviourNot documented
ChatGPT (OpenAI)OAI-SearchBot for search, GPTBot for trainingSettings are independent; a site opted out of OAI-SearchBot is not shown in ChatGPT search answers but can still appear as navigational linksIndex composition; how sources are weighted
PerplexityPerplexityBot for search; a page may also be visited when a user asks a questionPerplexityBot is designed to surface and link websites in search results; a page may be visited when a user asks a questionIndex composition; ranking criteria
Gemini (Google)Google Search infrastructure via the google_search toolThe model handles the entire workflow of searching, processing, and citing information automatically, generating and executing queries and grounding the answer in the resultsHow results are selected inside that workflow
Claude (Anthropic)Claude-SearchBot, which navigates the web to improve search result qualitySite owners can block the crawlers via robots.txtWhether a third-party index is involved; no provider is named

Two shortcuts to resist. The first assumes an engine simply mirrors one search index, so classic SEO would cover everything. The second assumes every engine runs a fully proprietary index. Neither is documented: the composition of the indexes and any third-party search agreements are not public, and we treat every claim about them as inference, not fact. Both shortcuts lead to wasted work: the first makes you optimize Google positions that one engine may never read, the second makes you ignore search fundamentals that still drive part of what the engines retrieve.

Three stages, routinely conflated

AI engines follow a retrieval-then-generate pipeline: they search, fetch pages, then write an answer grounded in what they fetched. Three stages have to succeed. First, crawling and indexing — can the engine's crawler fetch and store the page. Second, the search step — which index answered the internal query. Third, retrieval and citation — which fetched pages the model actually used and credited. Each stage has its own failure mode and its own diagnostic, and a symptom at the end — no citation — can come from any of the three.

This is why ranking first on Google is neither sufficient nor required. Not sufficient, because a stage-one block — OAI-SearchBot disallowed in robots.txt, for example — removes a page from ChatGPT search whatever its Google position, and that block is invisible in classic SEO tooling. Not required, because engines running their own crawlers can surface pages Google ranks modestly. In practice, the diagnosis order matters: confirm the crawler reaches you before rewriting a single page.

A page can also clear crawling and still fail at citation — in the reports we generate, this is the pattern we see most often. Retrieval and citation are separate steps, and published measurement shows the gap: Ahrefs' study of ChatGPT prompts, published in April 2026, found that 49.98% of the URLs retrieved during search — 23.4 million URLs — were cited in the final response, while 50.02% were not. Content can be fetched, read and used without being visibly credited; we call this the attribution gap, and it is one reason the difference between an AI mention and an AI citation matters when you measure. To detect stage-one activity directly, check your server logs for the documented crawler user agents: if OAI-SearchBot or PerplexityBot never appears there, the pipeline fails before any citation criterion applies.

Allow search without allowing training

The most actionable documented fact here is a robots.txt setting. OpenAI documents independent robots: OAI-SearchBot governs appearance in ChatGPT search, GPTBot governs training, and ChatGPT-User handles fetches a user triggers. The documentation, verified on 4 September 2026, states that each setting is independent of the others: a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot.

Perplexity and Anthropic document dedicated search crawlers as well — PerplexityBot and Claude-SearchBot — so review how each of them is handled, not only the OpenAI robots.

The common failure is precise: teams that blocked every OpenAI robot to opt out of training also opted out of search without meaning to, then spent weeks auditing content quality for a problem that lived in one file. The fix is a supported configuration, not a workaround:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

Two nuances. A site opted out of OAI-SearchBot can still appear as navigational links, per the same documentation — and a bare link is not a citation: the engine points at your domain without drawing on your content, which means your pages are not feeding the answer. And a robots.txt change is not picked up instantly; the crawlers re-read the file on a delay, so allow some time before judging the effect.

What the 2026 citation data measured

Beyond vendor documentation, published measurement tells us what the engines actually cite — as long as each figure travels with its scope.

Profound's AI Platform Citation Patterns research, published in June 2025 and based on a large multi-platform sample of citations, found that the platforms favour different source types: Wikipedia was ChatGPT's most cited source at 7.8% of total citations, Reddit led for Google AI Overviews at 2.2%, and .com domains carried over 80% of citations overall. The useful reading is less the exact percentages than the divergence: each platform has its own sourcing profile, so a source strategy tuned to one engine does not transfer as is. We cite the headline pattern here and link the full report in the sources.

Profound's FactCheck research, published in August 2026, which Profound describes as covering brand-related claims in AI answers, measured inaccuracy rates by answer engine ranging from 4.3% to 6.8%. The implication for sourcing: engines repeat what their retrieved sources say, including errors, so being the accurate, checkable source on your own subject is a defensive move, not just a visibility play. The scope matters: as Profound frames it, the claims measured concern brands rather than general knowledge, which is exactly the terrain where a brand's own documentation can serve as the correcting source.

The academic GEO paper by Aggarwal and co-authors, an arXiv preprint from November 2023, reports that generative engine optimization methods can boost visibility by up to 40% in generative engine responses. That figure is a maximum observed in the study, not an expected gain, and the paper itself states that efficacy varies across domains.

Across these studies, the mechanical pattern is consistent with the pipeline: the extraction step lifts passages, not pages. A section whose first sentences answer the question it opens with is easier to lift than an argument spread across the page — a writing decision, not a technical one.

On domain authority, the published data conflicts: some analyses find it correlates with citation frequency, others find close to no correlation. We do not present either camp as settled. When the evidence disagrees, the practitioner's move is to measure their own case — track whether the engines cite you on the prompts that matter — rather than optimize for a disputed signal. Our own reports are one way to run that measurement; a spreadsheet and a weekly round of prompts is another.

What we observe in Namedrop reports

What follows is qualitative observation from the reports we generate, last reviewed on 4 September 2026. We have no published sample or protocol yet, so no figures — these are patterns, stated as patterns, distinct from what the vendors document and from what the studies above measured.

  • Server-rendered pages clear the crawling stage more often than pages that need JavaScript to display their main content.
  • Pages with visible publication and update dates surface more on time-sensitive questions than undated pages.
  • Content signed by an identified author, and pages carrying original data, get cited more readily than anonymous aggregation.
  • Language changes the sources: on the French and English prompts we run for the same brand, the engines often select different sources, and a site visible in one language can be absent from answers in the other.

None of these observations carries a documented weight from any vendor; they are correlations we watch, and any of them may weaken as the engines change.

Where to start this week

Six steps, in order, each tied to a stage of the pipeline described above. None of them requires new tooling.

  1. Read your robots.txt and check how OAI-SearchBot, PerplexityBot and Claude-SearchBot are handled. If you blocked everything to opt out of training, allow the search crawlers explicitly (stage one).
  2. Load a key page with JavaScript disabled to approximate what a crawler sees. If the main content is missing, fix rendering before anything else (stage one).
  3. Show visible publication and update dates on the page, and keep them honest — a fresh date on unchanged content erodes trust when checked (retrieval).
  4. Sign content with a named author and a reference page for that author (retrieval and citation).
  5. Replace one generic claim with an original number from your own data — original figures are what answers quote (citation).
  6. Verify: search your server logs for the documented crawler user agents. Seeing OAI-SearchBot or PerplexityBot fetch your pages confirms the configuration works; their absence tells you where to look next (verification).

If a competitor keeps showing up in answers where you are absent, we have written about why competitors get cited by ChatGPT when you aren't, and about how to automate AI visibility checks once the basics are in place. Scan your site for free to see how each of the four engines handles you today, and which sources they cite instead.

Sources

Frequently asked questions

Why does ChatGPT never cite my website?
The most common cause we see is a stage-one block: OAI-SearchBot disallowed in robots.txt, often as a side effect of blocking GPTBot to opt out of training. OpenAI documents that a site opted out of OAI-SearchBot is not shown in ChatGPT search answers, whatever its Google rankings. Check your robots.txt and your server logs for the crawler before auditing content quality.
Can I get cited without letting models train on my content?
Yes. OpenAI documents independent settings for its crawlers: you can allow OAI-SearchBot, which governs appearance in ChatGPT search, while disallowing GPTBot, which governs training. Perplexity and Anthropic document separate search crawlers as well. Blocking every AI robot out of caution removes you from search results; the supported configuration is to allow the search crawlers and disallow the training ones.
Does domain authority still matter for getting cited by AI?
Published data conflicts. Some analyses find domain authority correlates with citation frequency; others find close to no correlation once other factors are accounted for. We do not treat either camp as settled. The practical answer is to measure your own case: track whether the engines cite you on the prompts that matter to your brand, and change one variable at a time.
Do AI engines prefer recently published content over older pages?
No vendor documents a freshness weight, so we treat this as observation rather than fact. In the reports we generate, pages with visible publication and update dates surface more often on time-sensitive questions, and undated pages rarely do; on evergreen questions the effect is weaker. Showing real dates and updating content when it actually changes is a low-cost signal with no documented downside.
Can one piece of content get cited by all four engines?
It happens, but nothing guarantees it. Each engine documents its own crawlers and applies its own selection behaviour, so the criteria differ. Stacking the signals that clear every pipeline — crawlable server-rendered pages, allowed search crawlers, visible dates, named authors, original data — improves the odds on all four engines at once without guaranteeing any single citation. Expect uneven results across engines and measure each one separately.