GPTBot: a training crawler run by OpenAI
GPTBot is a training crawler. It collects pages for this: future OpenAI models. Operator: OpenAI.
At a glance
- Operator
- OpenAI
- robots.txt token
GPTBot- Purpose
- Training. Training crawlers collect pages for future models. Keeping them out does not stop the crawlers that fetch pages for answers, which have names of their own.
- What it feeds
- future OpenAI models
What blocking it changes
If you keep GPTBot out, your pages are left out of what it collects for this: future OpenAI models. The crawlers that fetch pages for answers have names of their own, so they can still read your site.
User agent
The full user agent, from the operator’s documentation. A request from GPTBot carries this text, and a server log shows it.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbotrobots.txt lines for GPTBot
robots.txt is a plain text file at the root of a site (yoursite.com/robots.txt). A group starts with one or more User-agent lines that name crawlers, followed by Allow and Disallow lines for paths. A crawler named in a group follows the groups that name it and ignores the group for every crawler (User-agent: *).
To let it in
User-agent: GPTBot
Allow: /To keep it out
User-agent: GPTBot
Disallow: /Does your site let it in?
The free crawler check reads your robots.txt for GPTBot and requests your page with its user agent, so it shows whether a robots.txt rule or a firewall turns it away.