AI bot rules in robots.txt

Explicit rules for AI crawlers (GPTBot, ClaudeBot, Google-Extended and others) show that someone has decided who gets in.

Standard
Established

What it checks

The extension reads /robots.txt and looks for User-agent groups that name known AI crawlers. It then works out whether each named crawler can reach the site or is blocked from all of it.

The crawlers it recognises are GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI); ClaudeBot, Claude-Web, Claude-User and anthropic-ai (Anthropic); Google-Extended; GoogleOther; PerplexityBot and Perplexity-User; CCBot (Common Crawl); Bytespider (ByteDance); Amazonbot; Applebot-Extended; Meta-ExternalAgent; and cohere-ai.

This is the heaviest check in the extension. It is the only one that measures what a site has actually decided about AI agents. Naming crawlers only to block all of them does not earn a pass: the posture is what counts, not the naming.

A group “blocks everything” only when it disallows / without an Allow: / to override it. An empty Disallow: line allows everything.

Results

Status When
Pass At least one AI crawler is named, and at least one of the named crawlers can reach the site
Warn AI crawlers are named, but every one of them is fully blocked
Warn No AI crawler is named and nothing blocks them (only a * group, or no groups at all)
Fail No AI crawler is named and the * group disallows the whole site
Fail There is no robots.txt, so no rules can be expressed

How to fix

Add a User-agent group for each AI crawler you have an opinion about, stating what it may and may not read. For example, to allow search and answer engines while opting out of model training:

User-agent: OAI-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /

User-agent: *
Allow: /

If the whole site is blocked by User-agent: * / Disallow: /, check whether that is intended. It is often a staging robots.txt that shipped by mistake.

If you block every named AI crawler on purpose, the warning is expected. It is there to confirm the choice was deliberate.

Related: robots.txt agent-user policy covers the agents that fetch a page because a person asked, and Content Signals lets you say how content may be used rather than only whether it may be fetched.