Bytespider: a training crawler run by ByteDance
Bytespider is a training crawler. It collects pages for this: Doubao. Operator: ByteDance.
At a glance
- Operator
- ByteDance
- robots.txt token
Bytespider- Purpose
- Training. Training crawlers collect pages for future models. Keeping them out does not stop the crawlers that fetch pages for answers, which have names of their own.
- What it feeds
- Doubao
What blocking it changes
If you keep Bytespider out, your pages are left out of what it collects for this: Doubao. The crawlers that fetch pages for answers have names of their own, so they can still read your site.
User agent
The full user agent, from the operator’s documentation. A request from Bytespider carries this text, and a server log shows it.
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; [email protected])robots.txt lines for Bytespider
robots.txt is a plain text file at the root of a site (yoursite.com/robots.txt). A group starts with one or more User-agent lines that name crawlers, followed by Allow and Disallow lines for paths. A crawler named in a group follows the groups that name it and ignores the group for every crawler (User-agent: *).
To let it in
User-agent: Bytespider
Allow: /To keep it out
User-agent: Bytespider
Disallow: /Does your site let it in?
The free crawler check reads your robots.txt for Bytespider and requests your page with its user agent, so it shows whether a robots.txt rule or a firewall turns it away.