# GEO Audit Checklist: 12 Checks Before You Touch Your Content

Author: Quentin Megevand
Published: 2026-10-01
Source: https://getnamedrop.ai/en/blog/geo-audit-checklist-12-checks-before-you-touch-your-content

A GEO audit answers three questions before you edit anything: can AI engines reach your pages, can they understand your brand, and do they already cite you. The twelve checks below split into those three groups — access, content and entity, measurement — and each one ends with a pass or fail you record before changing a single page.

## What a GEO audit checks that an SEO audit doesn't

A GEO audit covers five areas: AI crawler access, structured data validity, extractable content structure, entity consistency, and a prompt-based citation baseline. Its output is a pass/fail list, not a to-do list: the point is to know what is broken before deciding what to rewrite.

The three groups: checks 1 to 4 cover access (robots.txt, rendering, server logs, llms.txt); checks 5 to 8 cover content and entity (answer structure, schema, fact consistency, freshness); checks 9 to 12 cover measurement (baseline, competitor share, accuracy, KPIs).

Parts of an SEO audit transfer directly: if a crawler cannot fetch a page, nothing downstream matters, and that logic holds for GPTBot as much as for Googlebot. Other checks have no SEO equivalent at all — no rank tracker tells you that ChatGPT states your old pricing. Rank and citation are also distinct outcomes: in a July 2025 analysis, Ahrefs reported that 14.40% of pages cited in AI Overviews do not rank in the SERPs, meaning below position 100. A page can rank without being cited, and be cited without ranking.

That is why a GEO audit measures its own signals beyond raw citation counts: mention rate, how often engines name the brand in answers; citation rate, how often they link one of your pages; and accuracy, whether what they say is true. These terms carry the measurement checks later in the article.

| | SEO audit | GEO audit |
| --- | --- | --- |
| What it measures | Rankings, index coverage, organic clicks | Mention rate, citation rate, answer accuracy |
| Main tooling | Search Console, rank trackers, site crawlers | Server logs, repeated prompt runs, AI performance reports |
| What a pass looks like | Page indexed and ranking for its query | Page reachable, facts consistent, brand cited in answers |

## Checks 1 to 4: can AI engines reach and read your pages

Start here; nothing else matters if crawlers cannot fetch the pages.

### Check 1: robots.txt rules per AI user agent

Read robots.txt line by line and record what each AI crawler may do: GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot and Google-Extended at minimum. Training crawlers and search crawlers are separate agents with separate consequences. OpenAI's crawler documentation, verified on 1 October 2026, separates GPTBot, which collects content for model training, from OAI-SearchBot, which surfaces sites in ChatGPT search; blocking one does not block the other. Anthropic's help center states that its bots respect do-not-crawl signals by honoring industry-standard directives in robots.txt. After any edit, re-test once the documented window has passed: OpenAI notes it can take roughly 24 hours from a robots.txt update for its search systems to adjust. Pass: every agent you intend to allow is explicitly allowed, and the training-versus-search decision is deliberate rather than a side effect of a wildcard rule.

### Check 2: pages render without JavaScript

Fetch each priority page with a plain curl request, or with JavaScript disabled in the browser, and read the raw HTML. Several AI crawlers do not execute JavaScript, so content injected client-side can be invisible to them even when the page looks complete to a visitor. Pass: the main content — the answer, the table, the product facts — is present in the initial HTML response.

### Check 3: crawler visits verified in server logs

An open robots.txt proves permission, not presence. Filter your access logs for the user-agent strings you allowed in check 1 and confirm recent hits on the pages that matter. Two caveats: any client can claim to be GPTBot, so the string alone is weak evidence, and some vendors publish official IP ranges that let you confirm a hit genuinely came from them. Pass: verified visits from the search and answer crawlers within a recent window. No visits at all is a finding to record and investigate, not a reason to start rewriting.

### Check 4: llms.txt, noted but not trusted

llms.txt is a proposed plain-text index of a site's key content for language models. Google's guide to optimizing for generative AI features, verified on 1 October 2026, is direct: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search." Record whether the file exists if the team wants one, but never weigh it as a pass/fail GEO factor. Pass here simply means the audit notes its presence or absence and moves on.

## Checks 5 to 8: is your content and entity data citation-ready

Run these four on a sample of priority pages; each takes under an hour.

### Check 5: answer-first structure on key pages

Open each priority page and read only the first screen. It should state the answer the page exists to give, not bury it mid-page under preamble. Then scan the rest: descriptive headings, lists and tables that carry facts in extractable form, instead of long undifferentiated prose. Pass: the core answer is readable without scrolling, and the supporting facts sit in structures a model can quote.

### Check 6: schema that matches visible content

Validate the structured data on the same sample: Organization for the site, Article for posts, FAQPage where questions are answered on the page. Parsing is not the bar. The markup must match what the page visibly says — same prices, same dates, same claims — because markup that diverges from visible content is a trust problem, not a technical one. Pass: valid markup and zero divergence between schema and page.

### Check 7: entity facts consistent everywhere

List the brand facts an engine might repeat — pricing, founding date, product claims, locations — and compare them across your site, LinkedIn, Crunchbase, G2 and the other third-party profiles that describe you. Engines answer from whichever source they retrieved, so one stale profile can outvote your homepage. On multilingual sites, extend the comparison across language versions: engines answer non-English prompts from whichever version they fetched, and a translated page with outdated pricing becomes the answer in that language. Pass: every inconsistency is flagged with its location; a clean first audit is rare.

### Check 8: authorship, dates and freshness signals

Check that priority pages show a named author, a visible publication date, and an updated date where content changed. These are verifiable trust signals rather than decoration, and the evidence for sourced content is specific: the GEO benchmark paper by Aggarwal et al., published in November 2023 and presented at KDD 2024, reports that optimization methods including adding statistics, quotations and source citations can boost visibility by up to 40% in generative engine responses. That figure is a maximum observed in their benchmark, not a gain to expect on any given page. Pass: dates and authorship visible on every priority page, and claims backed by cited sources where the page asserts facts.

## Checks 9 to 12: what AI engines currently say about you

### Check 9: a citation baseline from a locked prompt set

Write a fixed set of prompts your buyers would plausibly ask, lock it, and run it across ChatGPT, Perplexity, Gemini and Claude. Repeat each prompt enough times to smooth run-to-run variance, because single runs mislead, and record mention rate and citation rate per engine. A near-zero result is still a baseline; before changing anything, [diagnose why your brand gets zero mentions in ChatGPT](/en/blog/why-your-brand-gets-zero-mentions-in-chatgpt). Pass: a dated baseline table exists and can be reproduced.

### Check 10: competitor citation share on the same prompts

Re-read the same answers for competitors: which brands get mentioned instead of yours, and which domains the citations point to — their own sites, Reddit threads, review platforms, Wikipedia. The source mix shows where influence sits in your category. To keep the comparison stable over time, [benchmark your AI visibility against competitors with a locked prompt set](/en/blog/how-to-compare-your-ai-visibility-against-competitors) rather than with ad-hoc queries. Pass: a share-of-voice view per engine, with citing domains listed.

### Check 11: accuracy of what engines claim about your brand

Read every collected answer for factual errors: wrong pricing, discontinued products, invented features. When an engine states something false, the correction path is concrete. Fix the source pages the engine retrieved, which are often the inconsistent profiles found in check 7; use the platform's own feedback mechanisms; then re-run the prompt once the corrected content has been recrawled. Pass: every false claim logged with its suspected source and a correction owner.

### Check 12: KPIs defined and a measurement source chosen

Before the audit closes, decide what you will track and where the data comes from. Two first-party sources now exist at no cost. Google Search Console's generative AI performance report shows data about how your site performs in generative AI features on Google Search, per Google's help documentation consulted on 1 October 2026. Microsoft announced AI Performance in Bing Webmaster Tools in February 2026, described as insights that show how publisher content appears across Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations. On targets: no universal healthy citation rate exists, because the number depends on prompt set breadth and competitor density — a niche B2B tool and a consumer brand will see incomparable rates. Set an internal baseline from check 9 and measure movement against it rather than chasing an industry figure. Pass: KPIs named, reporting sources connected, baseline recorded.

## Where to start this week

Prioritize by dependency, not by effort-versus-impact scores: an access failure invalidates every downstream check, so it goes first regardless of how small the fix looks. Rewrite no content until checks 1 to 4 pass — edits on pages crawlers cannot reach change nothing.

A realistic first week: day one, run the robots.txt and rendering checks on your handful of most important pages; day two, do the entity fact comparison across your site and third-party profiles; days three to five, write the prompt set, run it across the four engines, and record the baseline.

Then set a cadence: re-run the measurement checks on a fixed schedule with the same prompt set so results stay comparable, and re-run the access checks after any migration or CMS change. Once the baseline exists, you can [automate the visibility checks instead of re-running them by hand](/en/blog/can-you-automate-ai-visibility-checks-here-s-how).

[Run a free Namedrop scan](/en?src=blog-article#scanner) to get the measurement checks — mentions, citations and competitor share — done for you before you start the manual part of the audit.

## Sources

- [OpenAI, Overview of OpenAI Crawlers, consulted 2026-10-01](https://developers.openai.com/api/docs/bots)
- [Google Search Central, Google's Guide to Optimizing for Generative AI Features on Google Search, consulted 2026-10-01](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide)
- [arXiv (KDD 2024 paper by Aggarwal et al.), GEO: Generative Engine Optimization, consulted 2026-10-01](https://arxiv.org/abs/2311.09735)
- [Google (Search Console Help), Generative AI performance report (Search), consulted 2026-10-01](https://support.google.com/webmasters/answer/16984139?hl=en)
- [Anthropic (Claude Help Center), Does Anthropic crawl data from the web, and how can site owners block the crawler?, consulted 2026-10-01](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler)
- [Microsoft Bing Webmaster Blog, Introducing AI Performance in Bing Webmaster Tools Public Preview, consulted 2026-10-01](https://blogs.bing.com/webmaster/2026/2/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview/)
- [Ahrefs, 76% of AI Overview Citations Pull From the Top 10, consulted 2026-10-01](https://ahrefs.com/blog/search-rankings-ai-citations/)

## Frequently asked questions

### What's the difference between an SEO audit and a GEO audit?

An SEO audit measures whether a site can rank: index coverage, rankings, organic traffic. A GEO audit measures whether AI engines such as ChatGPT, Perplexity, Gemini and Claude can reach your pages, understand your brand facts, and cite you in generated answers. Some checks overlap, crawlability above all, but the prompt-based citation baseline and the accuracy review of what engines say about you have no SEO equivalent. The two audits share pages, not pass criteria.

### Which AI crawlers should I allow in robots.txt?

Audit robots.txt for each agent separately — GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended — and distinguish training crawlers from search and answer crawlers, because blocking one does not block the other. OpenAI, for example, documents GPTBot for model training and OAI-SearchBot for ChatGPT search. If you want citations, allow the search and answer crawlers; whether to allow training crawlers is a separate policy decision to make deliberately, not by wildcard.

### What is an llms.txt file and does it help AI visibility?

llms.txt is a proposed plain-text file listing a site's key content for language models. Google states that you don't need to create new machine readable files, AI text files, markup or Markdown to appear in Google Search, and none of the vendor documentation cited in this article mentions using the file. Creating one is harmless, so note whether it exists, but never treat it as a pass or fail factor in a GEO audit.

### How do I verify GPTBot or PerplexityBot actually crawled my site?

Filter your server access logs for the crawler's user-agent string and look for recent hits on the pages you care about. Treat the string alone as weak evidence, because any client can claim to be GPTBot; some vendors publish official IP ranges so you can confirm a hit really came from them. No verified visits over a reasonable window, despite an open robots.txt, usually points to a discovery or rendering problem worth investigating.

### How often should you run a GEO audit?

Run the full twelve-check audit before any significant content work, and re-run the access checks after every migration, CMS change or robots.txt edit. The measurement checks follow a different rhythm: repeat the same locked prompt set on a fixed schedule, monthly for many teams, so results stay comparable between runs. Continuous monitoring can replace manual re-runs once a baseline exists, but it never replaces the structural checks.
