AI bot rules in robots.txt
Explicit rules for AI crawlers (GPTBot, ClaudeBot, Google-Extended and others) show that someone has decided who gets in.
- Category
- Bot access control
- Standard
- Established
What it checks
The extension reads /robots.txt and looks for User-agent groups that name
known AI crawlers. It then works out whether each named crawler can reach the
site or is blocked from all of it.
The crawlers it recognises are GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI); ClaudeBot, Claude-Web, Claude-User and anthropic-ai (Anthropic); Google-Extended; GoogleOther; PerplexityBot and Perplexity-User; CCBot (Common Crawl); Bytespider (ByteDance); Amazonbot; Applebot-Extended; Meta-ExternalAgent; and cohere-ai.
This is the heaviest check in the extension. It is the only one that measures what a site has actually decided about AI agents. Naming crawlers only to block all of them does not earn a pass: the posture is what counts, not the naming.
A group “blocks everything” only when it disallows / without an Allow: /
to override it. An empty Disallow: line allows everything.
Results
| Status | When |
|---|---|
| Pass | At least one AI crawler is named, and at least one of the named crawlers can reach the site |
| Warn | AI crawlers are named, but every one of them is fully blocked |
| Warn | No AI crawler is named and nothing blocks them (only a * group, or no groups at all) |
| Fail | No AI crawler is named and the * group disallows the whole site |
| Fail | There is no robots.txt, so no rules can be expressed |
How to fix
Add a User-agent group for each AI crawler you have an opinion about, stating
what it may and may not read. For example, to allow search and answer engines
while opting out of model training:
User-agent: OAI-SearchBot
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /
User-agent: *
Allow: /
If the whole site is blocked by User-agent: * / Disallow: /, check whether
that is intended. It is often a staging robots.txt that shipped by mistake.
If you block every named AI crawler on purpose, the warning is expected. It is there to confirm the choice was deliberate.
Related: robots.txt agent-user policy covers the agents that fetch a page because a person asked, and Content Signals lets you say how content may be used rather than only whether it may be fetched.