TL;DR: Free AI crawler robots.txt generator with a built-in knowledge base of 12 mainstream AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Bytespider...): purpose notes, allow/block guidance, and three-state decisions compiled into valid rules. Runs locally.
This tool runs entirely in your browser. Open DevTools → Network panel and search your input — it appears in no request.
Usage Guide
AI search (ChatGPT, Claude, Perplexity, Doubao) is absorbing the next generation of queries, and your robots.txt decides whether they can cite you. This tool ships a knowledge base of 12 mainstream AI crawlers — purpose, honest guidance, and a one-click decision UI that emits a valid robots.txt block. Runs entirely locally. The same decision framework produced this site’s own crawler policy; every recommendation below is the one we run.
1. Reading the three states
- Allow (Allow: /): full-site access — the default for public content sites; AI citations bring brand exposure and referral traffic.
- Block (Disallow: /): explicit refusal — right for paywalled content or crawlers whose data only feeds training with no citation flow back (e.g. Common Crawl).
- Follow (no group): the crawler falls through to your * wildcard — the honest "undecided" option, and a sensible default for training-only toggles like Google-Extended / Applebot-Extended.
2. Four field-tested calls
- Training crawlers vs. user crawlers: ChatGPT-User / Claude-User fetch pages live inside a user’s conversation — blocking breaks your site for AI users. Unless you have a hard reason, always allow.
- Block training-only crawlers with no citation flow (CCBot, meta-externalagent) at zero cost; allowing citation-driving crawlers (GPTBot, PerplexityBot) is free GEO referral traffic.
- Blocking Google-Extended does not affect Google Search indexing — it is a training toggle separate from Googlebot. Never cite "SEO risk" as the reason.
- Bytespider (ByteDance) crawls aggressively with a mixed compliance record; decide by crawl-cost. Small public sites lose nothing by blocking it.
3. After you generate
Merge the output into the robots.txt at your site root — if you already have rules, append the generated groups with a blank line between User-agent blocks. Validate with Google Search Console’s robots.txt tester. Vendors add UA tokens quarterly; re-run this tool each quarter — the crawler list and recommendations here are maintained on the same cadence.