AI crawler · ByteDance

Bytespider: what it is and how to block it

Bytespider is the crawler ByteDance documents for Toutiao Search, its Chinese search engine. ByteDance's own documentation doesn't mention AI training or say whether Bytespider follows robots.txt.

BrandVector Editorial · Last verified Oct 8, 2026

Operator
ByteDance
Type
Search
robots.txt token
Bytespider
Follows robots.txt
ByteDance doesn't say
IP ranges
Not published
Blocked by name
13% of top sites

What is Bytespider?

Bytespider is ByteDance's search crawler: it indexes pages for search results. In ByteDance's words:

“Toutiao Search's crawler user agent is "Bytespider", with a capital first letter.”

Original: 头条搜索的爬虫UA为“Bytespider”首写字母为大写

Bytespider user agent string

ByteDance documents these user agent strings for Bytespider:

Mozilla/5.0 (compatible; Bytespider; https://zhanzhang.toutiao.com/) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.0.0 Safari/537.36
Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; https://zhanzhang.toutiao.com/)
Mozilla/5.0 (iPhone; CPU iPhone OS 7_1_2 like Mac OS X) AppleWebKit/537.36 (KHTML, like Gecko) Version/7.0 Mobile Safari/537.36 (compatible; Bytespider; https://zhanzhang.toutiao.com/)

In robots.txt, use the token Bytespider, not the full string. Crawlers match on the token, and version numbers in the full string can change.

How many websites block Bytespider?

In our Oct 8, 2026 check of 817 of the most-visited websites, 13% block Bytespider by name for their whole site (105 sites).

robots.txt rule for BytespiderSitesShare
Blocks the whole site by name10513%
Blocks part of the site by name303.7%
Names it but doesn't block it40.5%
Blocked only by a catch-all (*) rule344.2%

For comparison, 0.1% of the same sites block Googlebot by name. The sites are the top 1,000 origins in the Chrome UX Report for August 2026; see how we measured.

How to block Bytespider in robots.txt

Add this group to the robots.txt file at the root of your domain to block Bytespider from your whole site:

User-agent: Bytespider
Disallow: /

To block only part of your site, list those paths instead:

User-agent: Bytespider
Disallow: /private/

A group that names Bytespider replaces your User-agent: * rules for it entirely (RFC 9309). If your catch-all group blocks paths that Bytespider should also stay out of, repeat them in its group. To let Bytespider in while a catch-all rule blocks others:

User-agent: *
Disallow: /

User-agent: Bytespider
Allow: /

ByteDance doesn't say whether Bytespider follows robots.txt, so a robots.txt block is a request it may not honor.

What happens if you block Bytespider?

ByteDance doesn't document what blocking Bytespider changes.

How to verify requests from Bytespider

“Bytespider's hostnames take the form *.bytedance.com. Anything that isn't *.bytedance.com is an impersonator.”

Original: Bytespider的hostname以*.bytedance.com的格式命名,非 *.bytedance.com即为冒充

See every AI crawler and how often top sites block it

Sources

  1. Toutiao Search Webmaster Platform: About Bytespider (in Chinese)
  2. RFC 9309: Robots Exclusion Protocol
  3. Chrome UX Report top 1,000 origins (global), August 2026

Rather have an expert do it?

Tell us about your site and goals, and we'll match you with an AI visibility (GEO) specialist. Free, no obligation.