Meta-ExternalAgent: a training crawler run by Meta
Meta-ExternalAgent is a training crawler. It collects pages for this: Meta AI. Operator: Meta.
At a glance
- Operator
- Meta
- robots.txt token
Meta-ExternalAgent- Purpose
- Training. Training crawlers collect pages for future models. Keeping them out does not stop the crawlers that fetch pages for answers, which have names of their own.
- What it feeds
- Meta AI
What blocking it changes
If you keep Meta-ExternalAgent out, your pages are left out of what it collects for this: Meta AI. The crawlers that fetch pages for answers have names of their own, so they can still read your site.
User agent
The full user agent, from the operator’s documentation. A request from Meta-ExternalAgent carries this text, and a server log shows it.
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)robots.txt lines for Meta-ExternalAgent
robots.txt is a plain text file at the root of a site (yoursite.com/robots.txt). A group starts with one or more User-agent lines that name crawlers, followed by Allow and Disallow lines for paths. A crawler named in a group follows the groups that name it and ignores the group for every crawler (User-agent: *).
To let it in
User-agent: Meta-ExternalAgent
Allow: /To keep it out
User-agent: Meta-ExternalAgent
Disallow: /Does your site let it in?
The free crawler check reads your robots.txt for Meta-ExternalAgent and requests your page with its user agent, so it shows whether a robots.txt rule or a firewall turns it away.