llms.txt generator

robots.txt rules for AI crawlers

The user-agent tokens AI companies document for their crawlers, each checked on the publisher's own page on 2026-10-07, and three ready-to-paste policies. Unlike llms.txt, these rules have a documented effect on crawlers that respect robots.txt.

AI crawler tokens by publisher

PublisherModel trainingAI search indexUser-requested fetch
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexity(none listed)PerplexityBotPerplexity-User
GoogleGoogle-Extended(Googlebot)
AppleApplebot-Extended(Applebot)
MetaMeta-ExternalAgentMeta-WebIndexerMeta-ExternalFetcher
AmazonAmazonbotAmzn-SearchBotAmzn-User
MistralMistralAI-TrainingMistralAI-IndexMistralAI-User
Common CrawlCCBot

Googlebot and Applebot are in parentheses: they are the regular search crawlers. Blocking them removes you from Google Search or Apple's search features, which is rarely what you want.

Details that change how you write the rules

Three ready-to-paste policies

Paste one block at the end of your robots.txt, after your existing rules. If your file already has groups for some of these crawlers, merge by hand. If your User-agent: * group disallows paths, add those same lines under the allowed crawlers (the generator reads your file and does it automatically).

1. Allow all AI crawlers

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: meta-webindexer
User-agent: meta-externalfetcher
User-agent: Amazonbot
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: MistralAI-Training
User-agent: MistralAI-Index
User-agent: MistralAI-User
Allow: /

2. Search yes, training no

AI search engines and user-requested fetches can read your pages (so you can be cited with a link); training crawlers and control tokens are disallowed.

# Allowed: AI search and user-requested fetches
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: meta-webindexer
User-agent: meta-externalfetcher
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: MistralAI-Index
User-agent: MistralAI-User
Allow: /

# Blocked: model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Amazonbot
User-agent: MistralAI-Training
Disallow: /

3. Block all AI crawlers

User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: meta-externalagent
User-agent: meta-webindexer
User-agent: meta-externalfetcher
User-agent: Amazonbot
User-agent: Amzn-SearchBot
User-agent: Amzn-User
User-agent: MistralAI-Training
User-agent: MistralAI-Index
User-agent: MistralAI-User
Disallow: /

Several User-agent lines followed by rules form one group (RFC 9309, section 2.1). Googlebot, Bingbot and other search engines are not affected by these blocks.

How to check your rules

  1. Open https://yoursite/robots.txt: it must answer 200 as plain text.
  2. Run the generator on your site: its report shows, for each of the 20 tokens, whether your current robots.txt allows the home page.