AI engines pull most of their citations from a handful of domain categories: community forums, encyclopedic references, news publishers, review and comparison platforms, video, and professional networks. Which category leads depends on the platform you measure, the industry your prompts cover, and how the study counts. This article explains all three variables.
The domain categories AI engines cite most
Every major citation study published since 2025 surfaces the same set of categories; what changes from study to study is the order inside that set. The table maps each category to the function it performs in an AI-generated answer and to the surface where it is strongest.
| Category | Role in AI answers | Representative domains | Where it is strongest |
|---|---|---|---|
| Community and forums | First-hand experience and opinions | Reddit, Quora, Stack Overflow | Product picks, troubleshooting, buying advice |
| Encyclopedic references | Definitions and entity facts | Wikipedia | Historically strongest on ChatGPT |
| News and professional publishers | Recency and expert reporting | Reuters, trade press | Time-sensitive and industry queries |
| Review and comparison platforms | Structured comparisons, ratings, pricing | G2, Capterra, Tripadvisor, NerdWallet | Commercial and B2B software queries |
| Video | Demonstrations and tutorials | YouTube | Google AI Overviews and AI Mode |
| Professional networks | People, companies, B2B expertise | Professional and B2B queries |
Community and forum content
Reddit is consistently among the single most-cited domains across AI platforms, and forums such as Quora and Stack Overflow follow the same logic: they hold first-hand experience, dissenting opinions and edge cases that no vendor page contains, the material an engine reaches for when the question is which option to buy or whether something actually works. That shared function also explains part of the volatility documented below: when a platform changes how much weight it gives forum content, the whole category moves at once.
Encyclopedic references
Wikipedia is a top-cited domain in effectively every published study and has historically been strongest on ChatGPT; Semrush's July to October 2025 measurements placed it among ChatGPT's most frequently cited domains before the sharp drop covered in the next section. Its function is entity grounding: definitions, dates and the factual skeleton of an answer.
News and professional publishers
News outlets and trade publications supply what forums and encyclopedias cannot: recency. Engines cite them on time-sensitive queries and on industry topics where a specialist publisher carries more depth than a generalist source, so their share of citations rises and falls with how newsy the prompt set is.
Review and comparison platforms
Review platforms hold structured comparisons, ratings and pricing, exactly the shape of data a recommendation answer needs. G2 and Capterra dominate B2B software queries, while Tripadvisor and NerdWallet play the same role in travel and finance, a pattern quantified in the industry section below.
Video and professional networks
On Google's AI surfaces, video leads outright. Ahrefs' analysis of the most-cited websites in Google AI Overviews, published in September 2026, found YouTube at the top, capturing 22.9% of citations. LinkedIn has risen into the top tier for professional and B2B queries: Profound reported in March 2026 that it had become the most-cited domain for professional queries across major platforms, its ChatGPT domain rank climbing from around #11 in November 2025 to around #5 by early 2026.
Category membership predicts citation likelihood better than raw popularity because an engine assembling an answer is filling functional slots — a definition, a first-hand account, a comparison, a recent fact — and it cites whichever domain fills each slot best for the query at hand. Popularity gets a domain crawled; function gets it cited. We cover the mechanism in more depth in how AI engines choose their sources.
How the top sources vary by AI platform
The ranking is platform-specific and volatile, with major shifts arriving within weeks. The clearest documented episode comes from Semrush's study of 230K prompts, run between July and October 2025. In that dataset, ChatGPT cited Reddit in close to 60% of prompt responses in early August 2025 before collapsing to around 10% by mid-September; Wikipedia dropped from appearing in roughly 55% of responses to less than 20% over the same window. No platform published an explanation for the shift.
At the category level, the engines lean differently. Google's AI Overviews and AI Mode draw heavily on video, with YouTube leading Ahrefs' September 2026 ranking. ChatGPT historically favoured encyclopedic and community sources, the two categories involved in the September 2025 drop. In Goodie's dataset covering ChatGPT and Gemini from October 2025 through March 2026, Wikipedia led the overall ranking, a sign the encyclopedic category holds up on Gemini as well. Perplexity's mix depends partly on which sites admit its crawler, a constraint that is mechanical rather than editorial.
One mechanical cause of divergence is crawler access. An engine can only cite content it can fetch, and site owners block crawlers selectively. Facebook's robots.txt, consulted on 24 September 2026, lists PerplexityBot as a user-agent with its own Disallow rules: paths Perplexity's crawler is told not to fetch, and therefore content Perplexity cannot draw on while an engine whose crawler faces no such rule can.
The practical consequence: a citation strategy tuned to one engine can lose most of its value in a single platform-side change, as September 2025 showed for anyone who had concentrated on Reddit visibility in ChatGPT. Spreading effort across categories — a community presence, a review-platform profile, strong owned content — is more durable than chasing whichever domain currently tops one engine's list.
Why the studies disagree on the numbers
Semrush measured Reddit in close to 60% of ChatGPT responses in August 2025; Goodie's October 2025 to March 2026 study has Wikipedia leading everything at 3.4%. Both figures are defensible, because they do not count the same thing.
The first method counts the share of responses that cite a domain at least once. A response citing Reddit alongside many other sources still counts fully for Reddit, so widely cited domains can appear in a majority of responses. Semrush's study works this way: it analyzed 230K prompts across ChatGPT search, Google AI Mode and Perplexity, taken as weekly snapshots between July 14 and October 12, 2025. Weekly snapshots also let a study show movement inside its own window, which is how the September 2025 drop was caught.
The second method computes a domain's share of all citations in the dataset. Goodie analyzed 58.6 million citations in ChatGPT and Gemini from October 2025 through March 2026; on that basis even the overall leader, Wikipedia, holds a 3.4% citation share, because the denominator is every citation to every domain. The same underlying behaviour can produce a large per-response percentage and a small share-of-citations percentage at the same time.
Prompt-set composition moves rankings further. A sample heavy in software queries lifts G2 and Capterra; one weighted toward travel lifts Tripadvisor; one built on professional queries lifts LinkedIn. And while each study discloses its dataset size, date range and platforms, none publishes a downloadable raw dataset. The percentages are vendor-reported, not independently audited, and not reproducible outside each vendor's tool: useful for direction, not for precision. That does not make them useless: independent studies converging on the same category-level pattern is meaningful even when the exact percentages differ.
Before relying on any citation ranking, apply a short test: what was counted, responses or citations; over which prompts; on which platforms. Honest studies that answer those differently will publish very different numbers for the same domain.
Industry changes the ranking more than the aggregate lists suggest
Inside a vertical, specialist domains routinely beat the famous generalist ones. Scrunch's industry breakdown of the most-cited AI sources, consulted on 24 September 2026, found that in finance NerdWallet takes the top spot while Reddit comes in at the bottom, and that Tripadvisor is far and away the leading citation source in travel and transportation, dominating hospitality and food as well. Peec AI's March 2026 analysis of 30 million sources across ChatGPT, Google AI Mode, Gemini, Perplexity and AI Overviews reports the same pattern in the hotel space, where Booking.com and Airbnb often outrank Reddit as AI sources.
This is the composition effect at work in aggregate leaderboards: an all-industries ranking is dominated by whichever verticals the prompt set over-samples. A dataset heavy in consumer software crowns review platforms; one heavy in travel crowns Tripadvisor. An aggregate list describes the study's prompt mix as much as the engines' preferences.
For a small or local business this is the useful finding. There is no realistic path onto Wikipedia, and earning durable Reddit presence or national news coverage is a long project. But every vertical has two or three domains that dominate its own citations, a review platform, an industry publication, a directory, a community, and presence there is attainable through a complete profile, current data and genuine reviews. Identifying those domains for your market is much of the answer to why competitors get cited by ChatGPT when you aren't.
Where to start this week
You do not need a published leaderboard to know which domains matter in your market; you can measure it directly.
First, write a couple of dozen prompts your buyers actually ask — comparisons, recommendations, best-in-town questions — and run them across at least two engines. Log every cited domain with the date of the run, because rankings move and undated data quickly loses its meaning.
Second, classify what you logged into the categories from the first section. Some domains you will never join; others — a review platform, a trade publication, a local directory, a community where your customers ask questions — you realistically can. Mark those.
Third, prioritise one vertical authority to build presence on and one owned-content improvement to ship, drawing on the content formats that earn competitors AI citations. Then re-run the same prompt set after a fixed interval and compare the two dated snapshots.
The deliverable is a citation map for your own niche: the domains your market's AI answers are actually built from, dated and refreshed on a schedule. It will not match any published aggregate ranking, and that is the point. As a labelled hypothetical, a regional accounting firm's map might contain a national comparison site, a professional association directory and an active practitioner forum — none glamorous, all citable, all attainable.
Run a free Namedrop scan to see which domains AI engines cite when they answer the questions your market asks.
Sources
- Ahrefs, The 50 Most-Cited Websites in Google AI Overviews (September 2026), consulted 2026-09-24
- Profound, LinkedIn is the Most-Cited Domain for Professional Queries in AI Search, consulted 2026-09-24
- Semrush, The Most-Cited Domains in AI: A 3-Month Study, consulted 2026-09-24
- Meta / Facebook, facebook.com/robots.txt, consulted 2026-09-24
- Goodie, Most Cited Domains in AI Search: Industry Breakdown, consulted 2026-09-24
- Scrunch, The Reddit paradox: What we learned from an industry breakdown of the most-cited AI sources, consulted 2026-09-24
- Peec AI, Top domains cited by AI search: Analysis based on 30M sources, consulted 2026-09-24