AI crawlers: what each bot does and what blocking it changes
The crawlers that read websites for AI assistants: who runs each one, what its work feeds, and the robots.txt lines to allow or block it.
AI crawlers are the bots that read websites for AI assistants. Search crawlers such as OAI-SearchBot build the index an assistant cites, user crawlers such as ChatGPT-User open a page when someone asks, and training crawlers such as GPTBot collect pages for future models. This list covers 22 of them, with robots.txt rules for each.
Three kinds of AI crawler
- Search
- Search crawlers read pages to build the index an assistant searches when it answers. A page they cannot read cannot be found or cited in those answers.
- User request
- User crawlers open one page when a person asks an assistant to read it. A page they cannot read cannot be opened in that conversation.
- Training
- Training crawlers collect pages for future models. Keeping them out does not stop the crawlers that fetch pages for answers, which have names of their own.
Every crawler, by what it feeds
Each name is the robots.txt token. Open one for its user agent and rules.
Search crawlers
| Crawler | Operator | What it feeds |
|---|---|---|
| OAI-SearchBot | OpenAI | ChatGPT search |
| Claude-SearchBot | Anthropic | Claude search |
| PerplexityBot | Perplexity | Perplexity |
| Googlebot | Google Search and AI Overviews | |
| Bingbot | Microsoft | Bing and Copilot |
| Applebot | Apple | Siri and Spotlight |
| DuckAssistBot | DuckDuckGo | DuckDuckGo answers |
| Amazonbot | Amazon | Alexa and Rufus |
| YouBot | You.com | You.com answers |
User crawlers
| Crawler | Operator | What it feeds |
|---|---|---|
| ChatGPT-User | OpenAI | ChatGPT |
| Claude-User | Anthropic | Claude |
| Perplexity-User | Perplexity | Perplexity |
| MistralAI-User | Mistral | Le Chat |
| Meta-ExternalFetcher | Meta | Meta AI |
Training crawlers
| Crawler | Operator | What it feeds |
|---|---|---|
| GPTBot | OpenAI | future OpenAI models |
| ClaudeBot | Anthropic | future Claude models |
| Google-Extended | Gemini training and grounding | |
| Applebot-Extended | Apple | Apple Intelligence training |
| Meta-ExternalAgent | Meta | Meta AI |
| CCBot | Common Crawl | the open dataset most models train on |
| Bytespider | ByteDance | Doubao |
| cohere-ai | Cohere | Cohere models |
A robots.txt template
robots.txt is a plain text file at the root of a site (yoursite.com/robots.txt). A group starts with one or more User-agent lines that name crawlers, followed by Allow and Disallow lines for paths. A crawler named in a group follows the groups that name it and ignores the group for every crawler (User-agent: *).
These lines let in every crawler that fetches pages for answers. The training crawlers follow behind # signs: letting them in is your choice.
Add these lines to the robots.txt file at the root of the site, and remove any group that disallows these crawlers. A crawler named in its own group follows only that group, so copy into it any Disallow lines it should keep.
# Crawlers that fetch pages for answers
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Googlebot
User-agent: Bingbot
User-agent: Applebot
User-agent: DuckAssistBot
User-agent: MistralAI-User
User-agent: Meta-ExternalFetcher
User-agent: Amazonbot
User-agent: YouBot
Allow: /
# Training crawlers: your choice. Remove the # signs to let them in.
# User-agent: GPTBot
# User-agent: ClaudeBot
# User-agent: Google-Extended
# User-agent: Applebot-Extended
# User-agent: Meta-ExternalAgent
# User-agent: CCBot
# User-agent: Bytespider
# User-agent: cohere-ai
# Allow: /Which of them does your site let in?
robots.txt is one way to turn a crawler away. A CDN or firewall setting is another, and the site still looks fine in a browser. The free crawler check reads your robots.txt and requests your page as each crawler, in a few seconds.
Questions about AI crawlers
Can AI assistants read your site?
The free crawler check answers in a few seconds, with no account and no email.