AI crawler directory

AI crawler bot directory

Current, source-graded facts and practical handling guidance for crawlers, live fetchers, agents, and control tokens.

93 sourced entries
59 operator sources
14 with live data
[01]

All Bots

Showing 93 of 93 sourced entries

Training crawler

Collects or controls access to content that can feed AI training data.

23 bots

AI2Bot

Training crawler

Allen Institute for AI crawler used to find web content for open language model datasets.

Allen Institute for AIHonors robots.txtOperator source
AI2Bot

Ai2Bot-Dolma

Training crawler

AI2 crawler token associated with Dolma/open language model dataset collection.

Allen Institute for AIHonors robots.txtOperator source
Ai2Bot-Dolma

anthropic-ai

Training crawler

Legacy Anthropic robots.txt token that predates the current ClaudeBot, Claude-User and Claude-SearchBot names.

AnthropicCompliance unverifiedTrusted crawler list
anthropic-ai

Applebot-Extended

Training crawler

Robots.txt control token for whether Applebot-crawled content may be used to train Apple foundation models.

AppleHonors robots.txtOperator source
Applebot-Extended

Bytespider

Training crawlerLive data

ByteDance crawler associated with training and powering AI products.

ByteDancePartial or user-triggeredTrusted crawler list
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Bytespider; [email protected]

CCBot

Training crawler

Common Crawl's crawler for building public web crawl datasets used by researchers and AI builders.

Common CrawlHonors robots.txtOperator source
CCBot/2.0 (https://commoncrawl.org/faq/)

ClaudeBot

Training crawlerLive data

Anthropic crawler for public web content that could contribute to Claude model training.

AnthropicHonors robots.txtOperator source
Mozilla/5.0 (compatible; ClaudeBot/1.0; [email protected])

cohere-training-data-crawler

Training crawler

Cohere training-data crawler token reported for downloading web data for enterprise language models.

CohereCompliance unverifiedPublic bot registry
cohere-training-data-crawler

DeepSeekBot

Training crawler

DeepSeek crawler token reported for training language models and improving AI products.

DeepSeekPartial or user-triggeredPublic bot registry
DeepSeekBot

Google-Extended

Training crawler

Robots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.

GoogleHonors robots.txtOperator source
Google-Extended

GPTBot

Training crawlerLive data

OpenAI crawler for content that may be used to improve generative AI foundation models.

OpenAIHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

GrokBot

Training crawler

xAI crawler token listed by public crawler directories for Grok-related crawling.

xAICompliance unverifiedTrusted crawler list
GrokBot

ICC-Crawler

Training crawler

NICT crawler for data used in artificial intelligence technologies and third-party research/commercial uses.

NICTHonors robots.txtPublic bot registry
ICC-Crawler

img2dataset

Training crawler

Open-source image dataset downloader token used to collect images for machine learning datasets.

img2datasetCompliance unverifiedOperator source
img2dataset

KimiBot

Training crawler

Moonshot AI crawler for content that may be used to improve its Kimi models.

Moonshot AIHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; KimiBot/1.0; +https://www.kimi.com/policies/kimi-crawlers

LAIONDownloader

Training crawler

LAION downloader token used in machine learning research dataset collection.

LAIONPartial or user-triggeredOperator source
LAIONDownloader

Meta-ExternalAgent

Training crawler

Meta crawler for indexing content directly for AI model training and product improvement use cases.

MetaHonors robots.txtOperator source
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)

MistralAI-Training

Training crawler

Mistral crawler for content that may be used to train its models, kept separate from its search index crawler.

Mistral AIHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)

PanguBot

Training crawler

Huawei crawler token reported for training data collection for the PanGu multimodal LLM.

HuaweiCompliance unverifiedPublic bot registry
PanguBot

SBIntuitionsBot

Training crawler

SB Intuitions crawler for data used in AI development and information analysis.

SB IntuitionsHonors robots.txtOperator source
SBIntuitionsBot

TerraCotta

Training crawler

Ceramic AI crawler token for downloading data used to train LLMs.

Ceramic AIHonors robots.txtOperator source
TerraCotta

VelenPublicWebCrawler

Training crawler

Velen crawler for business datasets and machine learning models.

VelenHonors robots.txtOperator source
VelenPublicWebCrawler

Webzio-Extended

Training crawler

Webz.io token covering whether crawled content may be included in the datasets it resells for AI and machine learning use.

Webz.ioHonors robots.txtOperator source
Webzio-Extended

AI search crawler

Indexes pages so AI search products can retrieve, rank, cite, or summarize them.

27 bots

amazon-kendra

AI search crawler

Amazon Kendra crawler token for intelligent enterprise search over configured content sources.

AmazonHonors robots.txtPublic bot registry
amazon-kendra

Amzn-SearchBot

AI search crawler

Amazon search crawler that indexes pages so they can be retrieved and cited in Amazon search and assistant answers.

AmazonHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36

Applebot

AI search crawler

Apple crawler for search experiences across Spotlight, Siri, Safari, and related Apple surfaces.

AppleHonors robots.txtOperator source
Applebot

atlassian-bot

AI search crawler

Atlassian Rovo crawler used to index connected website content for AI search, assistants, and agents.

AtlassianHonors robots.txtOperator source
atlassian-bot

Bingbot

AI search crawler

Microsoft Bing crawler used to crawl and index pages for Bing and Microsoft search-powered experiences.

MicrosoftHonors robots.txtTrusted crawler list
bingbot

Bravebot

AI search crawler

Brave Search crawler used to discover and index pages, including AI-search-adjacent retrieval.

BraveHonors robots.txtOperator source
Bravebot

Claude-SearchBot

AI search crawlerLive data

Anthropic search crawler that indexes content to improve Claude search result relevance and accuracy.

AnthropicHonors robots.txtOperator source
Mozilla/5.0 (compatible; Claude-SearchBot/1.0; [email protected])

Cloudflare-AutoRAG

AI search crawler

Cloudflare AutoRAG crawler used to index configured content for AI search applications.

CloudflareHonors robots.txtOperator source
Cloudflare-AutoRAG

ExaSearchBot

AI search crawler

Crawler for Exa's search index, which AI products and agents query to retrieve source pages.

ExaHonors robots.txtOperator source
Mozilla/5.0 (compatible; ExaSearchBot/1.0; +https://crawler.exa.ai/)

Googlebot

AI search crawler

Google Search crawler used to discover, crawl, render, and index pages for Google Search.

GoogleHonors robots.txtOperator source
Googlebot

iaskspider/2.0

AI search crawler

iAsk crawler used to provide answers to user queries.

iAskPartial or user-triggeredPublic bot registry
iaskspider/2.0

IbouBot

AI search crawler

Ibou crawler for building a graph representation of the web used in search.

IbouHonors robots.txtPublic bot registry
IbouBot

Kagibot

AI search crawler

Kagi search crawler that indexes pages for its search index and the answers built on it.

KagiHonors robots.txtOperator source
Mozilla/5.0 (compatible; Kagibot/1.0; +https://kagi.com/bot)

Kimi-SearchBot

AI search crawler

Moonshot AI search crawler that indexes pages so they can be retrieved and cited in Kimi answers.

Moonshot AIHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-SearchBot/1.0; +https://www.kimi.com/policies/kimi-crawlers

KlaviyoAIBot

AI search crawler

Klaviyo AI crawler for indexing configured content to tailor AI experiences and recommendations.

KlaviyoHonors robots.txtOperator source
KlaviyoAIBot

Meta-WebIndexer

AI search crawler

Meta crawler for improving Meta AI search result quality and source linking.

MetaCompliance unverifiedOperator source
meta-webindexer

MistralAI-Index

AI search crawler

Mistral automated crawler for indexing content used by Mistral AI search in Vibe.

Mistral AIHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)

OAI-SearchBot

AI search crawlerLive data

OpenAI search crawler for indexing pages that can appear in ChatGPT search results.

OpenAIHonors robots.txtOperator source
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot

PerplexityBot

AI search crawlerLive data

Perplexity crawler for surfacing and linking websites in Perplexity search results.

PerplexityHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

PetalBot

AI search crawler

Huawei crawler used for recommendations, assistant features, and AI search services.

HuaweiHonors robots.txtPublic bot registry
PetalBot

PhindBot

AI search crawler

Phind crawler token associated with AI-enhanced developer search.

PhindCompliance unverifiedPublic bot registry
PhindBot

ShapBot

AI search crawler

Parallel crawler for discovering and indexing websites for Parallel web APIs.

ParallelHonors robots.txtOperator source
ShapBot

Timpibot

AI search crawler

Timpi crawler reported for scraping data used in search and AI model training contexts.

TimpiCompliance unverifiedPublic bot registry
Timpibot

xAI-SearchBot

AI search crawler

Crawler token observed retrieving pages for xAI's Grok search and answer features.

xAICompliance unverifiedTrusted crawler list
Mozilla/5.0 (compatible; xAI-SearchBot/1.0; +https://x.ai)

YandexAdditional

AI search crawler

Yandex crawler token for data used in YandexGPT quick answers and additional analysis.

YandexHonors robots.txtOperator source
YandexAdditional

YandexAdditionalBot

AI search crawler

Yandex additional crawler token for YandexGPT-related answer and analysis features.

YandexHonors robots.txtOperator source
YandexAdditionalBot

YouBot

AI search crawlerLive data

You.com crawler for web search and AI answer experiences.

You.comHonors robots.txtOperator source
Mozilla/5.0 (compatible; YouBot (+http://www.you.com))

Live fetcher

Fetches pages because a user or agent asked for a specific URL or task.

20 bots

AmazonBuyForMe

Live fetcher

Amazon agent token reported for Buy for Me shopping actions directed by customers.

AmazonCompliance unverifiedPublic bot registry
AmazonBuyForMe

Amzn-User

Live fetcher

Amazon fetcher that retrieves a specific page because a person asked an Amazon assistant about it.

AmazonPartial or user-triggeredOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36

ChatGPT Agent

Live fetcher

OpenAI agent used when ChatGPT navigates websites for user-directed tasks.

OpenAICompliance unverifiedOperator source
ChatGPT Agent

ChatGPT-User

Live fetcherLive data

User-triggered OpenAI fetcher for ChatGPT and Custom GPT actions.

OpenAIPartial or user-triggeredOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot

Claude-Code

Live fetcher

Claude Code related agent token seen in crawler/user-agent lists.

AnthropicCompliance unverifiedPublic bot registry
Claude-Code

Claude-User

Live fetcher

Anthropic user-triggered fetcher for Claude answers that need a specific web page.

AnthropicHonors robots.txtOperator source
Claude-User

Claude-Web

Live fetcher

Reported Anthropic-related token seen in public crawler registries and Trakkr detection, but absent from Anthropic's current bot documentation.

AnthropicCompliance unverifiedPublic bot registry
Claude-Web

cohere-ai

Live fetcherLive data

Cohere token reported for retrieving data in response to user-initiated prompts.

CohereHonors robots.txtPublic bot registry
cohere-ai

DuckAssistBot

Live fetcher

DuckDuckGo AI assistant fetcher used by DuckAssist to retrieve content for real-time answers.

DuckDuckGoCompliance unverifiedTrusted crawler list
DuckAssistBot

Gemini-Deep-Research

Live fetcher

Gemini Deep Research agent token reported for collecting and scanning resources used in research answers.

GoogleCompliance unverifiedPublic bot registry
Gemini-Deep-Research

Google-Agent

Live fetcher

Google agent token listed by Cloudflare for AI assistant activity.

GoogleCompliance unverifiedTrusted crawler list
Google-Agent

Google-Gemini-CLI

Live fetcher

Gemini CLI related token listed in AI crawler registries for coding-agent activity.

GoogleCompliance unverifiedPublic bot registry
Google-Gemini-CLI

Google-GeminiNotebook

Live fetcher

Google fetcher that reads a URL because someone added it as a source inside Gemini or NotebookLM.

GooglePartial or user-triggeredOperator source
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)

Google-NotebookLM

Live fetcher

Superseded name for the Google fetcher that reads a URL when someone adds it as a source in Gemini or NotebookLM.

GoogleCompliance unverifiedPublic bot registry
Google-NotebookLM

GoogleAgent-Mariner

Live fetcher

Google AI agent token associated with browser-style task execution.

GoogleCompliance unverifiedPublic bot registry
GoogleAgent-Mariner

Kimi-User

Live fetcher

Moonshot AI fetcher that retrieves a specific page because a person asked Kimi about it.

Moonshot AIPartial or user-triggeredOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-User/1.0; +https://www.kimi.com/policies/kimi-crawlers

Meta-ExternalFetcher

Live fetcherLive data

Meta user-requested fetcher for AI and link features across Meta products.

MetaPartial or user-triggeredOperator source
meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)

MistralAI-User

Live fetcherLive data

Mistral user-action fetcher for Vibe responses that need a source page.

Mistral AIHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)

NovaAct

Live fetcher

Amazon Nova Act agent token reported for browser-style task execution.

AmazonCompliance unverifiedPublic bot registry
NovaAct

Perplexity-User

Live fetcherLive data

User-triggered Perplexity fetcher for pages needed to answer a specific question.

PerplexityPartial or user-triggeredOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

Other crawler

General-purpose, enterprise, or data-provider crawling with AI-adjacent use.

18 bots

Amazonbot

Other crawlerLive data

Amazon crawler used to improve its products and services and potentially train Amazon AI models.

AmazonHonors robots.txtOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36

ApifyBot

Other crawler

Token associated with crawlers run on the Apify scraping platform by its customers.

ApifyCompliance unverifiedTrusted crawler list
ApifyBot

bedrockbot

Other crawler

Amazon Bedrock web crawler connector token for customer-configured AI applications.

AmazonHonors robots.txtOperator source
bedrockbot

Brightbot

Other crawler

Bright Data's declared crawler for collecting public web data for its own datasets.

Bright DataCompliance unverifiedOperator source
Brightbot 1.0

Diffbot

Other crawlerLive data

Diffbot crawler for extracting structured web data and maintaining its knowledge graph.

DiffbotHonors robots.txtPublic bot registry
Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com)

FirecrawlAgent

Other crawler

Firecrawl agent token for AI scraping and web-to-LLM data extraction workflows.

FirecrawlHonors robots.txtPublic bot registry
FirecrawlAgent

Google-CloudVertexBot

Other crawler

Google crawler used for site-owner-requested crawls related to Vertex AI Agents.

GoogleHonors robots.txtOperator source
Google-CloudVertexBot

Google-Firebase

Other crawler

Google Firebase AI product token reported for app-related fetches.

GoogleCompliance unverifiedPublic bot registry
Google-Firebase

GoogleOther

Other crawler

Google generic crawler used by product teams for publicly accessible content fetches outside core Googlebot.

GoogleHonors robots.txtOperator source
GoogleOther

GoogleOther-Image

Other crawler

Google product-specific image crawler token for publicly accessible content fetches.

GoogleHonors robots.txtOperator source
GoogleOther-Image

GoogleOther-Video

Other crawler

Google product-specific video crawler token for public content fetches.

GoogleHonors robots.txtOperator source
GoogleOther-Video

ImagesiftBot

Other crawler

ImageSift crawler for public image and page data used in web intelligence products.

ImageSiftHonors robots.txtOperator source
ImagesiftBot

meta-externalads

Other crawler

Meta crawler documented alongside its other external agents, used for advertising related page reads.

MetaHonors robots.txtOperator source
meta-externalads/1.1

OAI-AdsBot

Other crawler

OpenAI crawler that reviews the safety and relevance of pages submitted as ChatGPT ad landing pages.

OpenAICompliance unverifiedOperator source
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot

omgili

Other crawler

Webz.io crawler for collecting web data sold through APIs and datasets.

Webz.ioHonors robots.txtOperator source
omgili

omgilibot

Other crawler

Legacy Omgili/Webz.io crawler token for web data collection.

Webz.ioHonors robots.txtPublic bot registry
omgilibot

Panscient

Other crawler

Panscient crawler for collecting and structuring business data with AI and machine learning.

PanscientHonors robots.txtOperator source
Panscient

Scrapy

Other crawler

Scrapy framework user-agent commonly used for web scraping, including AI and machine learning data extraction.

ZyteCompliance unverifiedPublic bot registry
Scrapy
93 source-graded pages
14 live-data cross-links
No invented telemetry
[02]

How to use this directory

Check the token

Use the exact user-agent token before writing a robots.txt rule or server policy.

Open the source

Each page labels whether its evidence comes from the operator, a trusted list, or a public registry.

Separate telemetry

Content pages explain the bot. Live crawl data stays on the data-site pages when Trakkr observes it.

See which AI crawlers actually visit your site

Track live AI crawler activity separately from this reference directory.

Get started free

14-day free trial · Cancel anytime