CCBot: a training crawler run by Common Crawl
CCBot is a training crawler. It collects pages for this: the open dataset most models train on. Operator: Common Crawl.
At a glance
- Operator
- Common Crawl
- robots.txt token
CCBot- Purpose
- Training. Training crawlers collect pages for future models. Keeping them out does not stop the crawlers that fetch pages for answers, which have names of their own.
- What it feeds
- the open dataset most models train on
What blocking it changes
If you keep CCBot out, your pages are left out of what it collects for this: the open dataset most models train on. The crawlers that fetch pages for answers have names of their own, so they can still read your site.
User agent
The full user agent, from the operator’s documentation. A request from CCBot carries this text, and a server log shows it.
CCBot/2.0 (https://commoncrawl.org/faq/)robots.txt lines for CCBot
robots.txt is a plain text file at the root of a site (yoursite.com/robots.txt). A group starts with one or more User-agent lines that name crawlers, followed by Allow and Disallow lines for paths. A crawler named in a group follows the groups that name it and ignores the group for every crawler (User-agent: *).
To let it in
User-agent: CCBot
Allow: /To keep it out
User-agent: CCBot
Disallow: /Does your site let it in?
The free crawler check reads your robots.txt for CCBot and requests your page with its user agent, so it shows whether a robots.txt rule or a firewall turns it away.