A GEO audit answers three questions before you edit anything: can AI engines reach your pages, can they understand your brand, and do they already cite you. The twelve checks below split into those three groups — access, content and entity, measurement — and each one ends with a pass or fail you record before changing a single page.
What a GEO audit checks that an SEO audit doesn't
A GEO audit covers five areas: AI crawler access, structured data validity, extractable content structure, entity consistency, and a prompt-based citation baseline. Its output is a pass/fail list, not a to-do list: the point is to know what is broken before deciding what to rewrite.
The three groups: checks 1 to 4 cover access (robots.txt, rendering, server logs, llms.txt); checks 5 to 8 cover content and entity (answer structure, schema, fact consistency, freshness); checks 9 to 12 cover measurement (baseline, competitor share, accuracy, KPIs).
Parts of an SEO audit transfer directly: if a crawler cannot fetch a page, nothing downstream matters, and that logic holds for GPTBot as much as for Googlebot. Other checks have no SEO equivalent at all — no rank tracker tells you that ChatGPT states your old pricing. Rank and citation are also distinct outcomes: in a July 2025 analysis, Ahrefs reported that 14.40% of pages cited in AI Overviews do not rank in the SERPs, meaning below position 100. A page can rank without being cited, and be cited without ranking.
That is why a GEO audit measures its own signals beyond raw citation counts: mention rate, how often engines name the brand in answers; citation rate, how often they link one of your pages; and accuracy, whether what they say is true. These terms carry the measurement checks later in the article.
| SEO audit | GEO audit | |
|---|---|---|
| What it measures | Rankings, index coverage, organic clicks | Mention rate, citation rate, answer accuracy |
| Main tooling | Search Console, rank trackers, site crawlers | Server logs, repeated prompt runs, AI performance reports |
| What a pass looks like | Page indexed and ranking for its query | Page reachable, facts consistent, brand cited in answers |
Checks 1 to 4: can AI engines reach and read your pages
Start here; nothing else matters if crawlers cannot fetch the pages.
Check 1: robots.txt rules per AI user agent
Read robots.txt line by line and record what each AI crawler may do: GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot and Google-Extended at minimum. Training crawlers and search crawlers are separate agents with separate consequences. OpenAI's crawler documentation, verified on 1 October 2026, separates GPTBot, which collects content for model training, from OAI-SearchBot, which surfaces sites in ChatGPT search; blocking one does not block the other. Anthropic's help center states that its bots respect do-not-crawl signals by honoring industry-standard directives in robots.txt. After any edit, re-test once the documented window has passed: OpenAI notes it can take roughly 24 hours from a robots.txt update for its search systems to adjust. Pass: every agent you intend to allow is explicitly allowed, and the training-versus-search decision is deliberate rather than a side effect of a wildcard rule.
Check 2: pages render without JavaScript
Fetch each priority page with a plain curl request, or with JavaScript disabled in the browser, and read the raw HTML. Several AI crawlers do not execute JavaScript, so content injected client-side can be invisible to them even when the page looks complete to a visitor. Pass: the main content — the answer, the table, the product facts — is present in the initial HTML response.
Check 3: crawler visits verified in server logs
An open robots.txt proves permission, not presence. Filter your access logs for the user-agent strings you allowed in check 1 and confirm recent hits on the pages that matter. Two caveats: any client can claim to be GPTBot, so the string alone is weak evidence, and some vendors publish official IP ranges that let you confirm a hit genuinely came from them. Pass: verified visits from the search and answer crawlers within a recent window. No visits at all is a finding to record and investigate, not a reason to start rewriting.
Check 4: llms.txt, noted but not trusted
llms.txt is a proposed plain-text index of a site's key content for language models. Google's guide to optimizing for generative AI features, verified on 1 October 2026, is direct: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search." Record whether the file exists if the team wants one, but never weigh it as a pass/fail GEO factor. Pass here simply means the audit notes its presence or absence and moves on.
Checks 5 to 8: is your content and entity data citation-ready
Run these four on a sample of priority pages; each takes under an hour.
Check 5: answer-first structure on key pages
Open each priority page and read only the first screen. It should state the answer the page exists to give, not bury it mid-page under preamble. Then scan the rest: descriptive headings, lists and tables that carry facts in extractable form, instead of long undifferentiated prose. Pass: the core answer is readable without scrolling, and the supporting facts sit in structures a model can quote.
Check 6: schema that matches visible content
Validate the structured data on the same sample: Organization for the site, Article for posts, FAQPage where questions are answered on the page. Parsing is not the bar. The markup must match what the page visibly says — same prices, same dates, same claims — because markup that diverges from visible content is a trust problem, not a technical one. Pass: valid markup and zero divergence between schema and page.
Check 7: entity facts consistent everywhere
List the brand facts an engine might repeat — pricing, founding date, product claims, locations — and compare them across your site, LinkedIn, Crunchbase, G2 and the other third-party profiles that describe you. Engines answer from whichever source they retrieved, so one stale profile can outvote your homepage. On multilingual sites, extend the comparison across language versions: engines answer non-English prompts from whichever version they fetched, and a translated page with outdated pricing becomes the answer in that language. Pass: every inconsistency is flagged with its location; a clean first audit is rare.
Check 8: authorship, dates and freshness signals
Check that priority pages show a named author, a visible publication date, and an updated date where content changed. These are verifiable trust signals rather than decoration, and the evidence for sourced content is specific: the GEO benchmark paper by Aggarwal et al., published in November 2023 and presented at KDD 2024, reports that optimization methods including adding statistics, quotations and source citations can boost visibility by up to 40% in generative engine responses. That figure is a maximum observed in their benchmark, not a gain to expect on any given page. Pass: dates and authorship visible on every priority page, and claims backed by cited sources where the page asserts facts.
Checks 9 to 12: what AI engines currently say about you
Check 9: a citation baseline from a locked prompt set
Write a fixed set of prompts your buyers would plausibly ask, lock it, and run it across ChatGPT, Perplexity, Gemini and Claude. Repeat each prompt enough times to smooth run-to-run variance, because single runs mislead, and record mention rate and citation rate per engine. A near-zero result is still a baseline; before changing anything, diagnose why your brand gets zero mentions in ChatGPT. Pass: a dated baseline table exists and can be reproduced.
Check 10: competitor citation share on the same prompts
Re-read the same answers for competitors: which brands get mentioned instead of yours, and which domains the citations point to — their own sites, Reddit threads, review platforms, Wikipedia. The source mix shows where influence sits in your category. To keep the comparison stable over time, benchmark your AI visibility against competitors with a locked prompt set rather than with ad-hoc queries. Pass: a share-of-voice view per engine, with citing domains listed.
Check 11: accuracy of what engines claim about your brand
Read every collected answer for factual errors: wrong pricing, discontinued products, invented features. When an engine states something false, the correction path is concrete. Fix the source pages the engine retrieved, which are often the inconsistent profiles found in check 7; use the platform's own feedback mechanisms; then re-run the prompt once the corrected content has been recrawled. Pass: every false claim logged with its suspected source and a correction owner.
Check 12: KPIs defined and a measurement source chosen
Before the audit closes, decide what you will track and where the data comes from. Two first-party sources now exist at no cost. Google Search Console's generative AI performance report shows data about how your site performs in generative AI features on Google Search, per Google's help documentation consulted on 1 October 2026. Microsoft announced AI Performance in Bing Webmaster Tools in February 2026, described as insights that show how publisher content appears across Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations. On targets: no universal healthy citation rate exists, because the number depends on prompt set breadth and competitor density — a niche B2B tool and a consumer brand will see incomparable rates. Set an internal baseline from check 9 and measure movement against it rather than chasing an industry figure. Pass: KPIs named, reporting sources connected, baseline recorded.
Where to start this week
Prioritize by dependency, not by effort-versus-impact scores: an access failure invalidates every downstream check, so it goes first regardless of how small the fix looks. Rewrite no content until checks 1 to 4 pass — edits on pages crawlers cannot reach change nothing.
A realistic first week: day one, run the robots.txt and rendering checks on your handful of most important pages; day two, do the entity fact comparison across your site and third-party profiles; days three to five, write the prompt set, run it across the four engines, and record the baseline.
Then set a cadence: re-run the measurement checks on a fixed schedule with the same prompt set so results stay comparable, and re-run the access checks after any migration or CMS change. Once the baseline exists, you can automate the visibility checks instead of re-running them by hand.
Run a free Namedrop scan to get the measurement checks — mentions, citations and competitor share — done for you before you start the manual part of the audit.
Sources
- OpenAI, Overview of OpenAI Crawlers, consulted 2026-10-01
- Google Search Central, Google's Guide to Optimizing for Generative AI Features on Google Search, consulted 2026-10-01
- arXiv (KDD 2024 paper by Aggarwal et al.), GEO: Generative Engine Optimization, consulted 2026-10-01
- Google (Search Console Help), Generative AI performance report (Search), consulted 2026-10-01
- Anthropic (Claude Help Center), Does Anthropic crawl data from the web, and how can site owners block the crawler?, consulted 2026-10-01
- Microsoft Bing Webmaster Blog, Introducing AI Performance in Bing Webmaster Tools Public Preview, consulted 2026-10-01
- Ahrefs, 76% of AI Overview Citations Pull From the Top 10, consulted 2026-10-01