ClaudeBot: a training crawler run by Anthropic
ClaudeBot is a training crawler. It collects pages for this: future Claude models. Operator: Anthropic.
At a glance
- Operator
- Anthropic
- robots.txt token
ClaudeBot- Purpose
- Training. Training crawlers collect pages for future models. Keeping them out does not stop the crawlers that fetch pages for answers, which have names of their own.
- What it feeds
- future Claude models
What blocking it changes
If you keep ClaudeBot out, your pages are left out of what it collects for this: future Claude models. The crawlers that fetch pages for answers have names of their own, so they can still read your site.
User agent
The full user agent, from the operator’s documentation. A request from ClaudeBot carries this text, and a server log shows it.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; [email protected])robots.txt lines for ClaudeBot
robots.txt is a plain text file at the root of a site (yoursite.com/robots.txt). A group starts with one or more User-agent lines that name crawlers, followed by Allow and Disallow lines for paths. A crawler named in a group follows the groups that name it and ignores the group for every crawler (User-agent: *).
To let it in
User-agent: ClaudeBot
Allow: /To keep it out
User-agent: ClaudeBot
Disallow: /Does your site let it in?
The free crawler check reads your robots.txt for ClaudeBot and requests your page with its user agent, so it shows whether a robots.txt rule or a firewall turns it away.