Google-Extended: a training crawler run by Google
Google-Extended is a training crawler. It collects pages for this: Gemini training and grounding. Operator: Google.
At a glance
- Operator
- robots.txt token
Google-Extended- Purpose
- Training. Training crawlers collect pages for future models. Keeping them out does not stop the crawlers that fetch pages for answers, which have names of their own.
- What it feeds
- Gemini training and grounding
What blocking it changes
If you keep Google-Extended out, your pages are left out of what it collects for this: Gemini training and grounding. The crawlers that fetch pages for answers have names of their own, so they can still read your site.
robots.txt lines for Google-Extended
robots.txt is a plain text file at the root of a site (yoursite.com/robots.txt). A group starts with one or more User-agent lines that name crawlers, followed by Allow and Disallow lines for paths. A crawler named in a group follows the groups that name it and ignores the group for every crawler (User-agent: *).
To let it in
User-agent: Google-Extended
Allow: /To keep it out
User-agent: Google-Extended
Disallow: /Does your site let it in?
Google-Extended is a name for robots.txt rules only. The crawler check reads robots.txt for it and sends no request as it.