Skip to content
Utilities2026-10-05Begin3 min read

Allow or Block AI Crawlers: GPTBot, ClaudeBot, PerplexityBot — a robots.txt Decision Guide

First, know exactly what you're allowing or blocking

AI companies run two classes of crawlers with different jobs — decide them separately:

CrawlerOwnerPurposeBlocking it affects
GPTBotOpenAIModel trainingNot ChatGPT Search inclusion
OAI-SearchBotOpenAIChatGPT Search indexingBlocks you out of ChatGPT Search
ChatGPT-UserOpenAIReal-time fetch on user requestUsers can't show your page to ChatGPT
ClaudeBotAnthropicModel trainingNot Claude search
Claude-Web / Claude-SearchBotAnthropicClaude search / real-timeSearch and live fetch
PerplexityBotPerplexitySearch indexingThe source pool for Perplexity answers
Google-ExtendedGoogleGemini trainingDoes not affect Google Search (separate from Googlebot)
GooglebotGoogleGoogle Search (incl. AI Overviews)Blocking it exits Google entirely
BytespiderByteDanceDoubao trainingHigh crawl volume — a load consideration
Applebot-ExtendedAppleApple Intelligence trainingNot Apple search
CCBotCommon CrawlOpen corpusIndirect training data for many models

The most common mistake: executing "don't train on my content" as a blanket block — and thereby also giving up the AI-search entry. Training and search are different UAs; you can block only the former.

The decision framework: three site types, three answers

Content sites / tool sites / blogs (our case) Allow search crawlers + allow training crawlers by default. Rationale: citation is distribution, AI search is a new traffic entry, and being in the training corpus means future models know you. The one exception: when load becomes a problem, rate-limit the heavy training crawlers (robots.txt Crawl-delay or CDN rate limiting).

Paid / membership content Block training crawlers (the content is the product), allow search crawlers but control snippet display with meta tags. Note robots.txt is a courtesy convention — the real defense for paid content is the auth wall itself.

Internal systems / backends Block everything (standard Disallow: /) — they shouldn't be crawlable at all.

Copy-paste robots.txt rules

Allow everything (content-site default):

User-agent: *
Allow: /

Block training only, keep search (cautious / paid):

# Block training crawlers
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

# Everything else (including search crawlers) allowed
User-agent: *
Allow: /

Rate-limit a heavy crawler (load consideration):

User-agent: Bytespider
Crawl-delay: 10
Allow: /

One rules note: the most specific UA section wins over the * section — so you can grant a single crawler an exception without touching the global policy.

After allowing: measure the payoff

Allowing isn't the finish line — you should see GEO returns:

  1. Watch the AI channel in analytics: GA4 and peers now itemize AI search as a source (this site sees a distinct "AI Assistant" channel) — the data section of the GEO introduction has real numbers
  2. Check server logs for crawler visits: grep for PerplexityBot / OAI-SearchBot fetch records after allowing — no crawling, no citation eligibility
  3. Just ask the AI: search your head keywords in ChatGPT/Perplexity and see whether the answers cite you — crudest and most effective check

Implementer's note

The real tradeoffs from configuring this site: our robots.txt is minimal (allow-all) because a tool site is openly distribution-oriented — but every UA string was verified against each vendor's official docs (case-sensitive; a typo equals no rule). Another field observation: Bytespider and YisouSpider crawl at volumes far beyond the rest (thousands of hits daily in our logs), so even when allowing them, CDN-level rate limiting is worth it for small servers — "allow" and "allow without limit" are different decisions.