Content Signals
Cloudflare Content Signals (Content-Signal in robots.txt) declare how content may be used for search, AI input, and AI training.
- Category
- Bot access control
- Standard
- Emerging checkEmerging
- Off by default, turn on in settings
What it checks
robots.txt says whether a crawler may fetch a page. Content Signals, proposed
by Cloudflare, say what the content may be used for once fetched. The extension
looks in /robots.txt for a Content-Signal: directive, or a comment that
mentions content-signal.
Results
| Status | When |
|---|---|
| Pass | A Content-Signal directive or comment is in robots.txt |
| N/A | No Content Signals are declared |
How to fix
Add a Content-Signal line to a User-agent group in robots.txt, setting
each use to yes or no:
User-agent: *
Content-Signal: search=yes, ai-input=yes, ai-train=no
Allow: /
search covers building a search index, ai-input covers using content to
answer a question in real time, and ai-train covers training or fine-tuning
models. Leaving a signal out expresses no preference for that use.