robots.txt rules for AI crawlers
The user-agent tokens AI companies document for their crawlers, each checked on the publisher's own page on 2026-10-07, and three ready-to-paste policies. Unlike llms.txt, these rules have a documented effect on crawlers that respect robots.txt.
AI crawler tokens by publisher
| Publisher | Model training | AI search index | User-requested fetch |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Perplexity | (none listed) | PerplexityBot | Perplexity-User |
Google-Extended | (Googlebot) | ||
| Apple | Applebot-Extended | (Applebot) | |
| Meta | Meta-ExternalAgent | Meta-WebIndexer | Meta-ExternalFetcher |
| Amazon | Amazonbot | Amzn-SearchBot | Amzn-User |
| Mistral | MistralAI-Training | MistralAI-Index | MistralAI-User |
| Common Crawl | CCBot |
Googlebot and Applebot are in parentheses: they are the regular search crawlers. Blocking them removes you from Google Search or Apple's search features, which is rarely what you want.
Details that change how you write the rules
Google-ExtendedandApplebot-Extendedare control tokens, not crawlers. Google says Google-Extended "does not impact a site's inclusion in Google Search"; Apple says Applebot-Extended "does not crawl webpages". Disallowing them opts you out of model training without leaving search.- User-requested fetchers may ignore robots.txt. Perplexity: "this fetcher generally ignores robots.txt rules". OpenAI on ChatGPT-User: "robots.txt rules may not apply". Meta and Amazon say the same for their user-triggered fetchers.
- A crawler with its own group stops reading
User-agent: *. RFC 9309 only falls back to*"if no matching group exists". AddUser-agent: GPTBot+Allow: /and GPTBot can now reach the/adminyou disallowed for everyone. Copy your*rules into the AI group: the generator does it for you. - Matching is case-insensitive (RFC 9309), so
meta-externalagentandMeta-ExternalAgentare the same token. - robots.txt is a request. Per the RFC, "these rules are not a form of access authorization". Well-behaved crawlers follow them; others don't.
- The list ages. New crawlers appear regularly. Re-check the publishers' pages linked above.
Three ready-to-paste policies
Paste one block at the end of your robots.txt, after your existing rules. If your file already has groups for some of these crawlers, merge by hand. If your User-agent: * group disallows paths, add those same lines under the allowed crawlers (the generator reads your file and does it automatically).
1. Allow all AI crawlers
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: meta-webindexer
User-agent: meta-externalfetcher
User-agent: Amazonbot
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: MistralAI-Training
User-agent: MistralAI-Index
User-agent: MistralAI-User
Allow: /
2. Search yes, training no
AI search engines and user-requested fetches can read your pages (so you can be cited with a link); training crawlers and control tokens are disallowed.
# Allowed: AI search and user-requested fetches
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: meta-webindexer
User-agent: meta-externalfetcher
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: MistralAI-Index
User-agent: MistralAI-User
Allow: /
# Blocked: model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: MistralAI-Training
Disallow: /
3. Block all AI crawlers
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: meta-webindexer
User-agent: meta-externalfetcher
User-agent: Amazonbot
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: MistralAI-Training
User-agent: MistralAI-Index
User-agent: MistralAI-User
Disallow: /
Several User-agent lines followed by rules form one group (RFC 9309, section 2.1). Googlebot, Bingbot and other search engines are not affected by these blocks.
How to check your rules
- Open
https://yoursite/robots.txt: it must answer 200 as plain text. - Run the generator on your site: its report shows, for each of the 20 tokens, whether your current robots.txt allows the home page.