Competitors get cited by AI engines because they publish content in formats those engines extract — direct answers, comparisons, fresh data — and on third-party surfaces engines trust, not because their brand is bigger. This article shows how to tell whether your gap is format, topic or surface, and how to close each one.
What AI engines are choosing when they cite a competitor's content
Start with a distinction worth drawing precisely: a competitor's page being cited as a source is not the same thing as the competitor's brand being recommended in the answer. A citation means the engine pulled facts from that URL; a recommendation means the answer names the brand as an option. The two have different causes — citations reward extractable, trustworthy pages, while recommendations reward how often and how favourably a brand appears across the sources an engine reads — and therefore different fixes. We cover the difference between an AI mention and an AI citation in a separate article.
The two can also split on the same page. When an engine cites a self-published best-tools listicle, the page owner collects the citation, but the brand the answer recommends may be a rival ranked inside that listicle. Publishing more content is therefore not automatically publishing content that works for you: a ranked listicle on your own blog can hand its citation to you and its recommendation to the competitors it names. We infer — no published study we have found quantifies brand recommendations by source type — that listicles, review pages and editorial coverage feed those recommendations more than vendors' own claims do.
Diagnosed at the page level, a citation gap resolves into three separate gaps. A format gap: the topic exists on your site, but the page is not structured for extraction. A topic gap: the question is answered by competitors and absent from your site. A surface gap: the citations go to pages nobody in the comparison owns — review platforms, community threads, reference sites — where competitors are present and you are not. Each gap has a different fix, which is why being told to build authority is not a diagnosis. If ChatGPT is your priority engine, we cover why competitors get cited by ChatGPT when you aren't in a dedicated article.
The content formats AI engines actually cite
Formats that get extracted
A short list of formats is built for extraction. Direct-answer explainers put the question in a heading and a complete answer in the first sentences under it. Comparison pages line up options in a table with checkable attributes. Original data pages publish numbers nobody else holds — we examine whether original data increases your odds of AI citation separately. Documentation and glossary-style definitions complete the set: both are dense in verifiable facts and stable in structure.
Structure is the prerequisite, not any single markup. Descriptive headings, an answer placed high on the page, tables and a high density of checkable facts make a passage easy to lift into a generated answer. No single trait guarantees extraction; a page can carry all of them and still lose to a stronger source.
Format work compounds existing SEO rather than replacing it. In Ahrefs' updated study of AI Overview citations, published on 2 March 2026, 37.9% of URLs cited in AI Overviews also appeared within the first 10 blocks of the search results. A page that already ranks is a page the engine has already selected once; restructure it before writing anything new.
How the format mix differs by engine
Engines do not cite the same mix of surfaces, so the same page performs differently across them. Profound's citation-patterns research, published on 5 June 2025, found that Reddit emerges as the leading source for both Google AI Overviews, at 2.2%, and Perplexity, at 6.6%. Semrush's three-month study of the most-cited domains in AI, published on 10 November 2025, found that ChatGPT cited Reddit in close to 60% of prompt responses in early August before collapsing to around 10% by mid-September. Source mixes are volatile, not fixed.
| Engine | Sourced citation pattern | Implication for your pages |
|---|---|---|
| ChatGPT | Reddit cited in close to 60% of prompt responses in early August 2025, around 10% by mid-September (Semrush, 10 November 2025) | Community discussion weighs heavily but swings; keep owned explainers as the stable base |
| Perplexity | Reddit is the leading cited source, at 6.6% (Profound, 5 June 2025) | Community threads plus direct-answer pages with visible facts |
| Google AI Overviews | Reddit leads at 2.2%; 37.9% of cited URLs also appear in the first 10 organic blocks (Profound, 5 June 2025; Ahrefs, 2 March 2026) | Pages that already rank, restructured for extraction |
| Gemini | No comparable public breakdown in the studies we cite | Apply the same structural rules; treat engine-specific claims as unverified |
Why third-party surfaces beat your own blog, and where the policy line sits
A large share of the citations competitors collect point to pages they do not own. The Reddit figures above concern a surface no brand controls, and review platforms, Wikipedia and media coverage play the same role. Third-party validation matters because a claim that exists only on your own domain is a claim only you make.
The legitimate route to those surfaces is unglamorous. Be reviewable: keep current profiles on the review platforms your buyers read, and ask real customers to write there. Contribute transparently to communities: answer questions under a disclosed affiliation, in threads where your product is genuinely relevant, and accept that some threads are not yours to win. Give journalists and analysts something citable: original data, a documented methodology, a number they can quote. Coverage follows citable material.
The shortcuts sit on the wrong side of documented policy. Undisclosed paid posting and bought or aged accounts conflict with Reddit's platform rules, and detection removes exactly the threads that were seeded. On the search side, Google's spam policies, dated 28 August 2026 in Google Search Central's documentation, define link spam as the practice of creating links to or from a site primarily for the purpose of manipulating search rankings; the same document's site reputation abuse policy covers content placed on third-party sites to exploit their ranking signals, which is where seeded affiliate placements sit. Wikipedia carries its own documented bar, notability, and treating it as a growth channel tends to end in reverted edits. Earn coverage that would survive an editor's scrutiny, or skip the surface.
Topic choices: the questions competitors answer that you don't
Format explains how competitors get extracted; topic explains where. The reliable way to find topic gaps is a fixed prompt set: write down the buying-relevant questions of your category and run each one in ChatGPT, Perplexity, Gemini and Google AI Overviews. For every prompt, record which competitor URLs are cited and note the format and topic of each cited page. One pass produces a map of the questions competitors answer that you do not.
Then tag each lost prompt with the gap it reveals, because each tag implies a different fix. Format gap: you cover the topic, but a competitor's better-structured page takes the citation — rework your page for extraction. Topic gap: the question has no page on your site — create one. Surface gap: the citations go to review platforms or community threads — earn presence there, because no owned page will win that prompt.
The tags give you the rewrite-versus-create rule: rewrite when the topic exists but the format fails extraction; create when the topic is absent from the site. An indexed, already-linked page is usually the faster page to restructure, which is consistent with the overlap between organic rankings and AI Overview citations noted above.
Freshness deserves careful phrasing. None of the vendor documentation we cite promises a recency multiplier, so we treat it as an inference: on questions whose answers age — prices, versions, statistics — a page whose facts are current is the harder one to replace. Maintain facts, not dates; bumping a last-updated date without changing the content is not maintenance, and it retains nothing.
What doesn't move the needle
Several popular tactics have no documented effect, and vendor documentation is the arbiter. Google Search Central's page on AI features and your website, dated 10 December 2025, states that you don't need to create new machine readable files, AI text files, or markup to appear in these features. For AI Overviews and AI Mode, nothing beyond content quality and normal indexing is required.
That settles the llms.txt question. The proposal — a plain-text file listing a site's key pages for language models — is widely recommended, while Google documents that its AI features need no such file. What other engines do with llms.txt is not documented by those engines; any claimed benefit there is inference, not finding. The file is cheap and harmless, so adding one costs little, but it belongs at the bottom of the list, not the top.
The same logic deprioritizes schema over-investment. Structured data has documented uses in search, but piling extra markup onto a page that lacks a direct answer, a table or a checkable fact fixes the wrong layer. Engines extract content; the fixes that pay are content fixes.
Where to start this week
The first deliverable is a one-page gap log, not a site overhaul.
- Write a fixed prompt set of a couple of dozen buying-relevant prompts — the questions a prospect in your category would type, in their words.
- Run the full set in ChatGPT, Perplexity, Gemini and Google AI Overviews, and log every cited competitor URL with the prompt it appeared on.
- Classify each cited page by format — explainer, comparison, data, documentation — by topic, and by surface: owned, review platform, community, reference.
- Tag each lost prompt as a format, topic or surface gap.
- Pick the two or three gaps with the clearest fix, schedule them, and leave the rest in the log.
Done by hand, expect the run and the logging to take between half a day and a full working day, and plan to repeat it monthly, because source mixes move: the swing in ChatGPT's Reddit citations that Semrush documented played out within weeks. Tooling changes the trade-off, not the work. Profound, one AI visibility tracking vendor, lists its Starter plan at $99/month on its public pricing page, consulted on 15 September 2026. That is a documented entry price for automating the run; the classification and the fixes remain yours. An agency adds the interpretation as well, at retainer prices too variable to quote responsibly. For a small team, the honest comparison is a day of manual work each month against a subscription: both produce the same gap log, and the tool buys frequency and consistency, not judgment.
Run a free Namedrop scan to see which prompts currently cite your competitors' content instead of yours.
Sources
- Profound, AI Platform Citation Patterns: How ChatGPT, Google AI Overviews, and Perplexity Source Information, consulted 2026-09-15
- Ahrefs, Update: 38% of AI Overview Citations Pull From The Top 10, consulted 2026-09-15
- Semrush, The Most-Cited Domains in AI: A 3-Month Study, consulted 2026-09-15
- Google Search Central, Spam Policies for Google Web Search, consulted 2026-09-15
- Google Search Central, AI Features and Your Website, consulted 2026-09-15
- Profound, Profound Pricing, consulted 2026-09-15