Allow or Block AI Crawlers: GPTBot, ClaudeBot, PerplexityBot — a robots.txt Decision Guide
First, know exactly what you're allowing or blocking
AI companies run two classes of crawlers with different jobs — decide them separately:
| Crawler | Owner | Purpose | Blocking it affects |
|---|---|---|---|
| GPTBot | OpenAI | Model training | Not ChatGPT Search inclusion |
| OAI-SearchBot | OpenAI | ChatGPT Search indexing | Blocks you out of ChatGPT Search |
| ChatGPT-User | OpenAI | Real-time fetch on user request | Users can't show your page to ChatGPT |
| ClaudeBot | Anthropic | Model training | Not Claude search |
| Claude-Web / Claude-SearchBot | Anthropic | Claude search / real-time | Search and live fetch |
| PerplexityBot | Perplexity | Search indexing | The source pool for Perplexity answers |
| Google-Extended | Gemini training | Does not affect Google Search (separate from Googlebot) | |
| Googlebot | Google Search (incl. AI Overviews) | Blocking it exits Google entirely | |
| Bytespider | ByteDance | Doubao training | High crawl volume — a load consideration |
| Applebot-Extended | Apple | Apple Intelligence training | Not Apple search |
| CCBot | Common Crawl | Open corpus | Indirect training data for many models |
The most common mistake: executing "don't train on my content" as a blanket block — and thereby also giving up the AI-search entry. Training and search are different UAs; you can block only the former.
The decision framework: three site types, three answers
Content sites / tool sites / blogs (our case) Allow search crawlers + allow training crawlers by default. Rationale: citation is distribution, AI search is a new traffic entry, and being in the training corpus means future models know you. The one exception: when load becomes a problem, rate-limit the heavy training crawlers (robots.txt Crawl-delay or CDN rate limiting).
Paid / membership content Block training crawlers (the content is the product), allow search crawlers but control snippet display with meta tags. Note robots.txt is a courtesy convention — the real defense for paid content is the auth wall itself.
Internal systems / backends
Block everything (standard Disallow: /) — they shouldn't be crawlable at all.
Copy-paste robots.txt rules
Allow everything (content-site default):
User-agent: *
Allow: /
Block training only, keep search (cautious / paid):
# Block training crawlers
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: /
# Everything else (including search crawlers) allowed
User-agent: *
Allow: /
Rate-limit a heavy crawler (load consideration):
User-agent: Bytespider
Crawl-delay: 10
Allow: /
One rules note: the most specific UA section wins over the * section — so you can grant a single crawler an exception without touching the global policy.
After allowing: measure the payoff
Allowing isn't the finish line — you should see GEO returns:
- Watch the AI channel in analytics: GA4 and peers now itemize AI search as a source (this site sees a distinct "AI Assistant" channel) — the data section of the GEO introduction has real numbers
- Check server logs for crawler visits: grep for PerplexityBot / OAI-SearchBot fetch records after allowing — no crawling, no citation eligibility
- Just ask the AI: search your head keywords in ChatGPT/Perplexity and see whether the answers cite you — crudest and most effective check
Implementer's note
The real tradeoffs from configuring this site: our robots.txt is minimal (allow-all) because a tool site is openly distribution-oriented — but every UA string was verified against each vendor's official docs (case-sensitive; a typo equals no rule). Another field observation: Bytespider and YisouSpider crawl at volumes far beyond the rest (thousands of hits daily in our logs), so even when allowing them, CDN-level rate limiting is worth it for small servers — "allow" and "allow without limit" are different decisions.
Related reading
Related Tools
Related Articles
What Is GEO? Generative Engine Optimization: The New Traffic Frontier of AI Search
GEO (Generative Engine Optimization) is to AI search what SEO is to Google — when ChatGPT, Perplexity, and Google's AI Overviews answer questions directly, users stop clicking links, and being cited by AI becomes the new traffic entry. This post breaks down the four essential differences from SEO, the four core practices (crawlability, structure, citability, fact density), and real GEO data from a working tool site.
llms.txt, Explained: A Site Manual for LLMs — and Whether Yours Deserves One
llms.txt is a 2024-proposed markdown file at your site root that tells large models what's worth reading. This post breaks down the spec structure (H1 + summary + tiered links), how it divides labor with robots.txt and sitemap, the honest controversies (Google's stance, adoption reality), and a real auto-generated implementation — with a skeleton you can copy.
The Complete Guide to PDF Tools: Merge, Split, Compress, and Convert — All in Your Browser
Everything you need to work with PDF files without uploading them: when to merge vs split, how PDF compression actually works, converting PDF to JPG and back, and why browser-local processing matters for contracts and ID documents.