---
title: "AI bot rules in robots.txt · Sitebulb Labs"
description: "Explicit rules for AI crawlers (GPTBot, ClaudeBot, Google-Extended and others) show that someone has decided who gets in."
url: https://labs.sitebulb.com/docs/checks/bot-access-control/ai-bot-rules/
---

# AI bot rules in robots.txt

Explicit rules for AI crawlers (GPTBot, ClaudeBot, Google-Extended and others) show that someone has decided who gets in.

- Category

  [Bot access control](https://labs.sitebulb.com/docs/checks/bot-access-control)

- Standard

  Established

## What it checks

The extension reads `/robots.txt` and looks for `User-agent` groups that name known AI crawlers. It then works out whether each named crawler can reach the site or is blocked from all of it.

The crawlers it recognises are GPTBot, OAI-SearchBot and ChatGPT-User (OpenAI); ClaudeBot, Claude-Web, Claude-User and anthropic-ai (Anthropic); Google-Extended; GoogleOther; PerplexityBot and Perplexity-User; CCBot (Common Crawl); Bytespider (ByteDance); Amazonbot; Applebot-Extended; Meta-ExternalAgent; and cohere-ai.

This is the heaviest check in the extension. It is the only one that measures what a site has actually decided about AI agents. Naming crawlers only to block all of them does not earn a pass: the posture is what counts, not the naming.

A group “blocks everything” only when it disallows `/` without an `Allow: /` to override it. An empty `Disallow:` line allows everything.

## Results

| Status   | When                                                                                        |
| -------- | ------------------------------------------------------------------------------------------- |
| **Pass** | At least one AI crawler is named, and at least one of the named crawlers can reach the site |
| **Warn** | AI crawlers are named, but every one of them is fully blocked                               |
| **Warn** | No AI crawler is named and nothing blocks them (only a `*` group, or no groups at all)      |
| **Fail** | No AI crawler is named and the `*` group disallows the whole site                           |
| **Fail** | There is no `robots.txt`, so no rules can be expressed                                      |

## How to fix

Add a `User-agent` group for each AI crawler you have an opinion about, stating what it may and may not read. For example, to allow search and answer engines while opting out of model training:

```txt
User-agent: OAI-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
User-agent: Google-Extended
User-agent: CCBot
Disallow: /

User-agent: *
Allow: /
```

If the whole site is blocked by `User-agent: *` / `Disallow: /`, check whether that is intended. It is often a staging `robots.txt` that shipped by mistake.

If you block every named AI crawler on purpose, the warning is expected. It is there to confirm the choice was deliberate.

Related: [robots.txt agent-user policy](https://labs.sitebulb.com/docs/checks/bot-access-control/robots-agent-user-policy) covers the agents that fetch a page because a person asked, and [Content Signals](https://labs.sitebulb.com/docs/checks/bot-access-control/content-signals) lets you say how content may be used rather than only whether it may be fetched.
