AI crawler · ByteDance
Bytespider: what it is and how to block it
Bytespider is the crawler ByteDance documents for Toutiao Search, its Chinese search engine. ByteDance's own documentation doesn't mention AI training or say whether Bytespider follows robots.txt.
BrandVector Editorial · Last verified Oct 8, 2026
- Operator
- ByteDance
- Type
- Search
- robots.txt token
Bytespider- Follows robots.txt
- ByteDance doesn't say
- IP ranges
- Not published
- Blocked by name
- 13% of top sites
What is Bytespider?
Bytespider is ByteDance's search crawler: it indexes pages for search results. In ByteDance's words:
“Toutiao Search's crawler user agent is "Bytespider", with a capital first letter.”
Original: 头条搜索的爬虫UA为“Bytespider”首写字母为大写
Bytespider user agent string
ByteDance documents these user agent strings for Bytespider:
Mozilla/5.0 (compatible; Bytespider; https://zhanzhang.toutiao.com/) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.0.0 Safari/537.36
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; https://zhanzhang.toutiao.com/)
Mozilla/5.0 (iPhone; CPU iPhone OS 7_1_2 like Mac OS X) AppleWebKit/537.36 (KHTML, like Gecko) Version/7.0 Mobile Safari/537.36 (compatible; Bytespider; https://zhanzhang.toutiao.com/)
In robots.txt, use the token Bytespider, not the full string. Crawlers match on the token, and version numbers in the full string can change.
How many websites block Bytespider?
In our Oct 8, 2026 check of 817 of the most-visited websites, 13% block Bytespider by name for their whole site (105 sites).
| robots.txt rule for Bytespider | Sites | Share |
|---|---|---|
| Blocks the whole site by name | 105 | 13% |
| Blocks part of the site by name | 30 | 3.7% |
| Names it but doesn't block it | 4 | 0.5% |
| Blocked only by a catch-all (*) rule | 34 | 4.2% |
For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.
How to block Bytespider in robots.txt
Add this group to the robots.txt file at the root of your domain to block Bytespider from your whole site:
User-agent: Bytespider Disallow: /
To block only part of your site, list those paths instead:
User-agent: Bytespider Disallow: /private/
A group that names Bytespider replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that Bytespider should also stay out of, repeat them in its group. To let Bytespider in while a catch-all rule blocks others:
User-agent: * Disallow: / User-agent: Bytespider Allow: /
ByteDance doesn't say whether Bytespider follows robots.txt, so a robots.txt block is a request it may not honor.
What happens if you block Bytespider?
ByteDance doesn't document what blocking Bytespider changes.
How to verify requests from Bytespider
“Bytespider's hostnames take the form *.bytedance.com. Anything that isn't *.bytedance.com is an impersonator.”
Original: Bytespider的hostname以*.bytedance.com的格式命名,非 *.bytedance.com即为冒充
See every AI crawler and how often top sites block it
Sources
Rather have an expert do it?
Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.