AI crawler bot directory
Current, source-graded facts and practical handling guidance for crawlers, live fetchers, agents, and control tokens.
All Bots
Training crawler
Collects or controls access to content that can feed AI training data.
AI2Bot
Training crawlerAllen Institute for AI crawler used to find web content for open language model datasets.
AI2BotAi2Bot-Dolma
Training crawlerAI2 crawler token associated with Dolma/open language model dataset collection.
Ai2Bot-Dolmaanthropic-ai
Training crawlerLegacy Anthropic robots.txt token that predates the current ClaudeBot, Claude-User and Claude-SearchBot names.
anthropic-aiApplebot-Extended
Training crawlerRobots.txt control token for whether Applebot-crawled content may be used to train Apple foundation models.
Applebot-ExtendedBytespider
Training crawlerLive dataByteDance crawler associated with training and powering AI products.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Bytespider; [email protected]CCBot
Training crawlerCommon Crawl's crawler for building public web crawl datasets used by researchers and AI builders.
CCBot/2.0 (https://commoncrawl.org/faq/)ClaudeBot
Training crawlerLive dataAnthropic crawler for public web content that could contribute to Claude model training.
Mozilla/5.0 (compatible; ClaudeBot/1.0; [email protected])cohere-training-data-crawler
Training crawlerCohere training-data crawler token reported for downloading web data for enterprise language models.
cohere-training-data-crawlerDeepSeekBot
Training crawlerDeepSeek crawler token reported for training language models and improving AI products.
DeepSeekBotGoogle-Extended
Training crawlerRobots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.
Google-ExtendedGPTBot
Training crawlerLive dataOpenAI crawler for content that may be used to improve generative AI foundation models.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbotGrokBot
Training crawlerxAI crawler token listed by public crawler directories for Grok-related crawling.
GrokBotICC-Crawler
Training crawlerNICT crawler for data used in artificial intelligence technologies and third-party research/commercial uses.
ICC-Crawlerimg2dataset
Training crawlerOpen-source image dataset downloader token used to collect images for machine learning datasets.
img2datasetKimiBot
Training crawlerMoonshot AI crawler for content that may be used to improve its Kimi models.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; KimiBot/1.0; +https://www.kimi.com/policies/kimi-crawlersLAIONDownloader
Training crawlerLAION downloader token used in machine learning research dataset collection.
LAIONDownloaderMeta-ExternalAgent
Training crawlerMeta crawler for indexing content directly for AI model training and product improvement use cases.
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)MistralAI-Training
Training crawlerMistral crawler for content that may be used to train its models, kept separate from its search index crawler.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)PanguBot
Training crawlerHuawei crawler token reported for training data collection for the PanGu multimodal LLM.
PanguBotSBIntuitionsBot
Training crawlerSB Intuitions crawler for data used in AI development and information analysis.
SBIntuitionsBotTerraCotta
Training crawlerCeramic AI crawler token for downloading data used to train LLMs.
TerraCottaVelenPublicWebCrawler
Training crawlerVelen crawler for business datasets and machine learning models.
VelenPublicWebCrawlerWebzio-Extended
Training crawlerWebz.io token covering whether crawled content may be included in the datasets it resells for AI and machine learning use.
Webzio-ExtendedAI search crawler
Indexes pages so AI search products can retrieve, rank, cite, or summarize them.
amazon-kendra
AI search crawlerAmazon Kendra crawler token for intelligent enterprise search over configured content sources.
amazon-kendraAmzn-SearchBot
AI search crawlerAmazon search crawler that indexes pages so they can be retrieved and cited in Amazon search and assistant answers.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36Applebot
AI search crawlerApple crawler for search experiences across Spotlight, Siri, Safari, and related Apple surfaces.
Applebotatlassian-bot
AI search crawlerAtlassian Rovo crawler used to index connected website content for AI search, assistants, and agents.
atlassian-botBingbot
AI search crawlerMicrosoft Bing crawler used to crawl and index pages for Bing and Microsoft search-powered experiences.
bingbotBravebot
AI search crawlerBrave Search crawler used to discover and index pages, including AI-search-adjacent retrieval.
BravebotClaude-SearchBot
AI search crawlerLive dataAnthropic search crawler that indexes content to improve Claude search result relevance and accuracy.
Mozilla/5.0 (compatible; Claude-SearchBot/1.0; [email protected])Cloudflare-AutoRAG
AI search crawlerCloudflare AutoRAG crawler used to index configured content for AI search applications.
Cloudflare-AutoRAGExaSearchBot
AI search crawlerCrawler for Exa's search index, which AI products and agents query to retrieve source pages.
Mozilla/5.0 (compatible; ExaSearchBot/1.0; +https://crawler.exa.ai/)Googlebot
AI search crawlerGoogle Search crawler used to discover, crawl, render, and index pages for Google Search.
Googlebotiaskspider/2.0
AI search crawleriAsk crawler used to provide answers to user queries.
iaskspider/2.0IbouBot
AI search crawlerIbou crawler for building a graph representation of the web used in search.
IbouBotKagibot
AI search crawlerKagi search crawler that indexes pages for its search index and the answers built on it.
Mozilla/5.0 (compatible; Kagibot/1.0; +https://kagi.com/bot)Kimi-SearchBot
AI search crawlerMoonshot AI search crawler that indexes pages so they can be retrieved and cited in Kimi answers.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-SearchBot/1.0; +https://www.kimi.com/policies/kimi-crawlersKlaviyoAIBot
AI search crawlerKlaviyo AI crawler for indexing configured content to tailor AI experiences and recommendations.
KlaviyoAIBotMeta-WebIndexer
AI search crawlerMeta crawler for improving Meta AI search result quality and source linking.
meta-webindexerMistralAI-Index
AI search crawlerMistral automated crawler for indexing content used by Mistral AI search in Vibe.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)OAI-SearchBot
AI search crawlerLive dataOpenAI search crawler for indexing pages that can appear in ChatGPT search results.
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbotPerplexityBot
AI search crawlerLive dataPerplexity crawler for surfacing and linking websites in Perplexity search results.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)PetalBot
AI search crawlerHuawei crawler used for recommendations, assistant features, and AI search services.
PetalBotPhindBot
AI search crawlerPhind crawler token associated with AI-enhanced developer search.
PhindBotShapBot
AI search crawlerParallel crawler for discovering and indexing websites for Parallel web APIs.
ShapBotTimpibot
AI search crawlerTimpi crawler reported for scraping data used in search and AI model training contexts.
TimpibotxAI-SearchBot
AI search crawlerCrawler token observed retrieving pages for xAI's Grok search and answer features.
Mozilla/5.0 (compatible; xAI-SearchBot/1.0; +https://x.ai)YandexAdditional
AI search crawlerYandex crawler token for data used in YandexGPT quick answers and additional analysis.
YandexAdditionalYandexAdditionalBot
AI search crawlerYandex additional crawler token for YandexGPT-related answer and analysis features.
YandexAdditionalBotYouBot
AI search crawlerLive dataYou.com crawler for web search and AI answer experiences.
Mozilla/5.0 (compatible; YouBot (+http://www.you.com))Live fetcher
Fetches pages because a user or agent asked for a specific URL or task.
AmazonBuyForMe
Live fetcherAmazon agent token reported for Buy for Me shopping actions directed by customers.
AmazonBuyForMeAmzn-User
Live fetcherAmazon fetcher that retrieves a specific page because a person asked an Amazon assistant about it.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36ChatGPT Agent
Live fetcherOpenAI agent used when ChatGPT navigates websites for user-directed tasks.
ChatGPT AgentChatGPT-User
Live fetcherLive dataUser-triggered OpenAI fetcher for ChatGPT and Custom GPT actions.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/botClaude-Code
Live fetcherClaude Code related agent token seen in crawler/user-agent lists.
Claude-CodeClaude-User
Live fetcherAnthropic user-triggered fetcher for Claude answers that need a specific web page.
Claude-UserClaude-Web
Live fetcherReported Anthropic-related token seen in public crawler registries and Trakkr detection, but absent from Anthropic's current bot documentation.
Claude-Webcohere-ai
Live fetcherLive dataCohere token reported for retrieving data in response to user-initiated prompts.
cohere-aiDuckAssistBot
Live fetcherDuckDuckGo AI assistant fetcher used by DuckAssist to retrieve content for real-time answers.
DuckAssistBotGemini-Deep-Research
Live fetcherGemini Deep Research agent token reported for collecting and scanning resources used in research answers.
Gemini-Deep-ResearchGoogle-Agent
Live fetcherGoogle agent token listed by Cloudflare for AI assistant activity.
Google-AgentGoogle-Gemini-CLI
Live fetcherGemini CLI related token listed in AI crawler registries for coding-agent activity.
Google-Gemini-CLIGoogle-GeminiNotebook
Live fetcherGoogle fetcher that reads a URL because someone added it as a source inside Gemini or NotebookLM.
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)Google-NotebookLM
Live fetcherSuperseded name for the Google fetcher that reads a URL when someone adds it as a source in Gemini or NotebookLM.
Google-NotebookLMGoogleAgent-Mariner
Live fetcherGoogle AI agent token associated with browser-style task execution.
GoogleAgent-MarinerKimi-User
Live fetcherMoonshot AI fetcher that retrieves a specific page because a person asked Kimi about it.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-User/1.0; +https://www.kimi.com/policies/kimi-crawlersMeta-ExternalFetcher
Live fetcherLive dataMeta user-requested fetcher for AI and link features across Meta products.
meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)MistralAI-User
Live fetcherLive dataMistral user-action fetcher for Vibe responses that need a source page.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)NovaAct
Live fetcherAmazon Nova Act agent token reported for browser-style task execution.
NovaActPerplexity-User
Live fetcherLive dataUser-triggered Perplexity fetcher for pages needed to answer a specific question.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)SEO tool crawler
Checks pages for SEO, writing, or content tooling powered by AI.
Social crawler
Reads pages for social previews, sharing, or social-platform AI features.
FacebookBot
Social crawlerMeta crawler historically documented for Facebook crawling and AI-related training uses.
FacebookBotfacebookexternalhit
Social crawlerMeta link preview crawler used when content is shared on Meta family apps.
facebookexternalhitTikTokSpider
Social crawlerByteDance/TikTok crawler token reported alongside AI crawling lists.
TikTokSpiderOther crawler
General-purpose, enterprise, or data-provider crawling with AI-adjacent use.
Amazonbot
Other crawlerLive dataAmazon crawler used to improve its products and services and potentially train Amazon AI models.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36ApifyBot
Other crawlerToken associated with crawlers run on the Apify scraping platform by its customers.
ApifyBotbedrockbot
Other crawlerAmazon Bedrock web crawler connector token for customer-configured AI applications.
bedrockbotBrightbot
Other crawlerBright Data's declared crawler for collecting public web data for its own datasets.
Brightbot 1.0Diffbot
Other crawlerLive dataDiffbot crawler for extracting structured web data and maintaining its knowledge graph.
Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com)FirecrawlAgent
Other crawlerFirecrawl agent token for AI scraping and web-to-LLM data extraction workflows.
FirecrawlAgentGoogle-CloudVertexBot
Other crawlerGoogle crawler used for site-owner-requested crawls related to Vertex AI Agents.
Google-CloudVertexBotGoogle-Firebase
Other crawlerGoogle Firebase AI product token reported for app-related fetches.
Google-FirebaseGoogleOther
Other crawlerGoogle generic crawler used by product teams for publicly accessible content fetches outside core Googlebot.
GoogleOtherGoogleOther-Image
Other crawlerGoogle product-specific image crawler token for publicly accessible content fetches.
GoogleOther-ImageGoogleOther-Video
Other crawlerGoogle product-specific video crawler token for public content fetches.
GoogleOther-VideoImagesiftBot
Other crawlerImageSift crawler for public image and page data used in web intelligence products.
ImagesiftBotmeta-externalads
Other crawlerMeta crawler documented alongside its other external agents, used for advertising related page reads.
meta-externalads/1.1OAI-AdsBot
Other crawlerOpenAI crawler that reviews the safety and relevance of pages submitted as ChatGPT ad landing pages.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbotomgili
Other crawlerWebz.io crawler for collecting web data sold through APIs and datasets.
omgiliomgilibot
Other crawlerLegacy Omgili/Webz.io crawler token for web data collection.
omgilibotPanscient
Other crawlerPanscient crawler for collecting and structuring business data with AI and machine learning.
PanscientScrapy
Other crawlerScrapy framework user-agent commonly used for web scraping, including AI and machine learning data extraction.
ScrapyHow to use this directory
Check the token
Use the exact user-agent token before writing a robots.txt rule or server policy.
Open the source
Each page labels whether its evidence comes from the operator, a trusted list, or a public registry.
Separate telemetry
Content pages explain the bot. Live crawl data stays on the data-site pages when Trakkr observes it.
See which AI crawlers actually visit your site
Track live AI crawler activity separately from this reference directory.
Get started free14-day free trial · Cancel anytime