What is Brightbot? AI crawler guide

Brightbot by Bright Data: Bright Data's declared crawler for collecting public web data for its own datasets. Check its reported user-agent, robots.txt behavior, source, and verification guidance.

Bright Data's declared crawler for collecting public web data for its own datasets.

What is Brightbot?

Brightbot is a web crawler operated by Bright Data that collects publicly available web data for the company's own datasets. It identifies itself with the user-agent token Brightbot and follows the rules set in robots.txt. Bright Data also offers a separate proxy network that routes customer traffic through residential IP addresses, and that traffic does not identify itself as Brightbot. The company encourages site owners to use its collectors.txt file in addition to robots.txt, and it reviews those rules before they take effect.

What it's for

If you run a website, Brightbot may crawl your pages to gather data that Bright Data uses in its commercial datasets. This could mean your public content appears in datasets sold to third parties. Blocking Brightbot in robots.txt can stop this declared crawler, but it does not affect traffic coming through Bright Data's residential proxy network, which does not identify itself.

How to handle Brightbot

To block Brightbot, add a disallow rule for the user-agent token Brightbot in your robots.txt file. Bright Data also recommends using its collectors.txt file to set additional rules, which it reviews before applying. Keep in mind that this only affects the declared Brightbot crawler and not other traffic from Bright Data's proxy services.

robots.txt rule

User-agent: Brightbot Disallow: /

Blocking cost

Blocking Brightbot may prevent your content from being included in Bright Data's datasets, which could reduce your visibility in data products that rely on those datasets.

Examples

Related bots

Frequently Asked Questions

What is Brightbot?

Brightbot is a web crawler from Bright Data that collects public web data for the company's own datasets. It identifies itself with the user-agent token Brightbot.

How can I block Brightbot?

You can block Brightbot by adding a disallow rule for the user-agent token Brightbot in your robots.txt file. Bright Data also suggests using its collectors.txt file for additional rules.

Does blocking Brightbot stop all Bright Data traffic?

No, blocking Brightbot only stops the declared crawler. Bright Data's residential proxy network routes customer traffic through residential IP addresses that do not identify themselves, so that traffic is not affected by a Brightbot block.

What is collectors.txt?

Collectors.txt is a file that Bright Data recommends site owners use to set crawling rules. Bright Data reviews these rules before they take effect, offering an additional control beyond robots.txt.

What happens if I don't block Brightbot?

If you do not block Brightbot, it may crawl your site and include your public data in Bright Data's commercial datasets, which could be sold to third parties.

Data & Sources