What is cohere-training-data-crawler? AI crawler guide

cohere-training-data-crawler by Cohere: Cohere training-data crawler token reported for downloading web data for enterprise language models. Check its reported user-agent, robots.txt behavior, source, and verification guidance.

Cohere training-data crawler token reported for downloading web data for enterprise language models.

What is cohere-training-data-crawler?

cohere-training-data-crawler is a web crawler operated by Cohere that downloads publicly available web data to train enterprise language models. It identifies itself with the user-agent token cohere-training-data-crawler. The crawler's purpose is to gather text from the open web for use in training Cohere's AI systems. Its behavior and compliance with robots.txt are unverified, meaning there is no confirmed documentation on how it respects standard exclusion rules. Site owners can use the token in their robots.txt file to signal an opt-out from this specific crawling activity.

What it's for

For a site owner, this crawler represents Cohere's effort to collect training data for its language models. If you do not want your content used in this way, you can block the crawler. Allowing it may contribute your site's public text to Cohere's training datasets, which could influence how their models perform on topics related to your content.

cohere-training-data-crawler collects pages for model training. Training inclusion is not the same as being cited, so measure where AI answers actually cite your site before drawing conclusions from crawl logs.

How to handle cohere-training-data-crawler

To prevent cohere-training-data-crawler from accessing your site, add a robots.txt rule targeting its user-agent token. The page will display the exact snippet you need. Because the crawler's robots.txt posture is unverified, there is no guarantee it will honor the rule, but it is the standard method to express your opt-out preference.

robots.txt rule

User-agent: cohere-training-data-crawler Disallow: /

Blocking cost

Blocking cohere-training-data-crawler may prevent your content from being included in Cohere's training data, which could reduce your site's indirect influence on their language models.

Examples

Related bots

Frequently Asked Questions

What does cohere-training-data-crawler do?

It crawls websites to download publicly available text for training Cohere's enterprise language models.

How can I stop cohere-training-data-crawler from crawling my site?

You can add a robots.txt rule that disallows the user-agent token cohere-training-data-crawler. The page provides the exact snippet.

Does cohere-training-data-crawler obey robots.txt?

Its compliance is unverified. There is no official documentation confirming that it respects robots.txt rules, so blocking it may not be effective.

What happens if I allow cohere-training-data-crawler?

Your site's public text may be downloaded and used to train Cohere's language models, potentially affecting how those models generate content related to your domain.

Is cohere-training-data-crawler associated with any other bots?

Based on available information, there are no related bots explicitly linked to this crawler.

Data & Sources