# Which AI Crawlers to Allow in robots.txt, and Which to Block

Author: Quentin Megevand
Published: 2026-10-10
Source: https://getnamedrop.ai/en/blog/which-ai-crawlers-should-you-allow-in-robots-txt

For most brands that want AI engines to cite them, allow the search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) in robots.txt and treat training controls (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended) as a separate licensing decision. Remember that robots.txt is voluntary, and some user-triggered agents, such as Perplexity-User, may not follow it.

## AI crawlers fall into three groups: search, training and user-triggered

robots.txt is a plain-text file at the root of each host telling crawlers which paths they may fetch. It asks; it blocks nothing technically. Search crawlers fetch pages so engines can cite them, training crawlers collect content for models, and user-triggered agents fetch a page a person asked an assistant to read.

Each agent is set separately. OpenAI's documentation, verified on 10 October 2026, says a webmaster "can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot."

### The user agents each vendor documents, in one table

| Vendor | User agent | Purpose | Respects robots.txt (vendor docs) | Verified on |
|---|---|---|---|---|
| OpenAI | OAI-SearchBot | Search | Yes, independent of GPTBot | 2026-10-10 |
| OpenAI | GPTBot | Training | Yes, independent of OAI-SearchBot | 2026-10-10 |
| OpenAI | ChatGPT-User | User-triggered | May not apply to user-initiated requests | 2026-10-10 |
| Anthropic | Claude-SearchBot | Search | Yes, all Anthropic bots | 2026-10-10 |
| Anthropic | ClaudeBot | Training | Yes | 2026-10-10 |
| Anthropic | Claude-User | User-triggered | Yes | 2026-10-10 |
| Perplexity | PerplexityBot | Search | Not stated | 2026-10-10 |
| Perplexity | Perplexity-User | User-triggered | Generally ignores it | 2026-10-10 |
| Google | Google-Extended | Training | Yes; no effect on Search | 2026-10-10 |
| Apple | Applebot-Extended | Training | Yes, documented opt-out | 2026-10-10 |

Google-Extended and Applebot-Extended are opt-out tokens named in robots.txt, not separate crawlers that fetch pages.

## Allow AI search, decide training separately: the robots.txt rules for each goal

A crawler follows the most specific User-agent group naming it and ignores `*`. Adding a named group for one bot removes the rules it read under `*`, so repeat any it still needs. Within a group, the Robots Exclusion Protocol says "The most specific match found MUST be used": the longest path wins.

Opting out of training is a licensing and data-use choice; OpenAI documents search and training settings as independent.

### Appear in AI search, opt out of training

```
# Search: allowed, private path repeated
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Disallow: /admin/

# Training: opted out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

User-agent: *
Disallow: /admin/
```

### Block all AI crawlers

```
User-agent: OAI-SearchBot
User-agent: GPTBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: ClaudeBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
```

Googlebot is untouched, so Google Search is unaffected.

### Allow everything

```
User-agent: *
Allow: /
```

This also permits training use.

## What robots.txt cannot do for AI crawlers

Compliance is voluntary and cannot be enforced. Perplexity's documentation says its user-triggered fetcher "generally ignores robots.txt rules". A blocked URL can still appear as a bare link when other sites link to it. To keep content out of results, use noindex, X-Robots-Tag or authentication; a crawler blocked by robots.txt never sees noindex.

### Google-Extended and AI Overviews

Google's documentation, verified on 10 October 2026, says Google-Extended does not affect inclusion or ranking in Google Search. It does not link the token to AI Overviews, a Search feature, so we infer Googlebot governs them.

### Content already crawled

None of the vendor text we verified on 10 October 2026 says a new Disallow removes content already collected.

### Subdomains, CDNs and staging hosts

Each host, including subdomains, staging sites and CDN hostnames, needs its own robots.txt. Cloudflare's documentation, verified on 10 October 2026, says its managed setting "will prepend our managed `robots.txt` before your existing `robots.txt`".

## How to check that AI crawlers follow your rules in server logs

A user-agent string proves nothing: anyone can send it.

- Filter logs for the documented user agents.
- Match source IPs against published ranges, such as openai.com/gptbot.json and openai.com/searchbot.json.
- Use reverse DNS where a vendor documents it.
- Compare fetched paths with your Disallow rules after crawlers reread the file.

Disallowed fetches from verified IPs mean non-compliance; from other IPs, spoofing.

## Does allowing AI crawlers increase citations

Allowing search crawlers is a precondition for citation, not a guarantee. OpenAI links allowing OAI-SearchBot to appearing in search results, but no source we used measures a citation gain. Blocking keeps control and removes eligibility; allowing gives eligibility, nothing more. Being cited then depends on [how AI engines select the sources they cite](/en/blog/how-ai-engines-choose-sources).

## Where to start this week

- Fetch the live robots.txt on every host.
- List which AI agents have named groups and which fall back to `*`.
- Choose one of the three goals with the client, apply its template, then continue the [GEO audit checklist, starting with crawler access](/en/blog/geo-audit-checklist-12-checks-before-you-touch-your-content).
- After a few days, check logs against the IP files.
- Recheck vendor docs each quarter, because agent names change.

Once your robots.txt reflects the goal you chose, [run a free Namedrop scan](/en?src=blog-article#scanner) to see which engines currently cite your brand and which competitors they cite instead.

## Sources

- [OpenAI, Overview of OpenAI Crawlers, consulted 2026-10-10](https://developers.openai.com/api/docs/bots)
- [Anthropic (Claude Help Center), Does Anthropic crawl data from the web, and how can site owners block the crawler?, consulted 2026-10-10](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler)
- [Perplexity, Perplexity Crawlers, consulted 2026-10-10](https://docs.perplexity.ai/guides/bots)
- [Apple, About Applebot - Apple Support, consulted 2026-10-10](https://support.apple.com/en-us/119829)
- [IETF (RFC Editor), RFC 9309: Robots Exclusion Protocol, consulted 2026-10-10](https://datatracker.ietf.org/doc/rfc9309/)
- [Google, Google's common crawlers, consulted 2026-10-10](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers)
- [Cloudflare, robots.txt setting - Cloudflare bot solutions docs, consulted 2026-10-10](https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/)

## Frequently asked questions

### Does blocking Google-Extended remove my site from AI Overviews?

Google's crawler documentation, verified on 10 October 2026, states that Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal. The text we verified does not link the token to AI Overviews. Because AI Overviews are a Search feature, our inference is that Googlebot access, not Google-Extended, governs them. Check Google's page before you change either rule.

### Is robots.txt legally binding for AI companies?

No source in this article settles that. robots.txt is a voluntary protocol defined by the IETF: it states your preferences, compliant crawlers follow them, and nothing technically stops others from ignoring them. Any legal effect depends on jurisdiction, contracts and terms of service, and this answer is not legal advice. For content that must stay private, use authentication rather than robots.txt.

### Can I see AI crawler visits in my server logs?

Yes. Filter your access logs for documented user agents such as GPTBot, OAI-SearchBot or PerplexityBot. Because anyone can send a user-agent string, match each source IP against the ranges the vendor publishes, such as openai.com/gptbot.json and openai.com/searchbot.json. Disallowed fetches from verified IPs show non-compliance, and the same name from other IPs shows spoofing.

### Does adding a Disallow rule remove content AI companies already crawled?

Not according to the documentation we checked. The vendor pages we verified on 10 October 2026 explain how to stop future crawling, and none of that text says a new Disallow rule deletes content collected earlier. Treat a Disallow as applying from the next time a crawler reads your file, and contact the vendor directly if you need data already gathered removed.

### Does llms.txt replace robots.txt for AI crawlers?

No. llms.txt is a proposed convention, not an IETF standard, and it does not control crawling: it suggests content for language models to read. Rules that allow or block a crawler still live in robots.txt, under the Robots Exclusion Protocol. A site can publish both files, but only robots.txt carries the access rules that compliant crawlers check.
