See the request. Check the citation. Measure the visit.
Trakkr records AI crawler and fetcher activity by page, then compares it with separately observed citations, AI referral traffic, access signals and next actions. You get a useful funnel without pretending a crawl caused an answer.
Crawler monitoring evidence model
What is recorded?
A matching HTTP request, normalized page, time, method, response status, referrer when present, crawler class, user-agent signal, source and collection method.
What is verified?
The UI distinguishes a signature match from stronger network or provider evidence. User-agent alone never proves who sent a request.
What is joined?
Crawler events, citations, AI referrals and optional search data are compared on normalized page URLs and aligned windows, with their source and limits intact.
Separate purpose before policy
Training crawlers
GPTBot, ClaudeBotFetch pages that may be used to improve models. Their robots controls are separate from search crawler controls where the operator documents separate tokens.
Search and indexing crawlers
OAI-SearchBot, Claude-SearchBot, PerplexityBotBuild or refresh retrieval indexes used by AI search products. A successful request can support access or eligibility, but does not predict ranking or citation.
User-requested fetchers
ChatGPT-User, Claude-User, Perplexity-UserRetrieve a page in response to a user action. A request does not reveal the conversation, prove the page was read, or prove that an answer cited it.
Control tokens
Google-ExtendedSet a robots.txt policy but may not appear as a distinct HTTP user-agent. Google says Google-Extended does not affect Search inclusion or ranking.
The crawler to citation to referral funnel
Request reaches a layer
Edge, CDN, WAF or origin records a matching request. Cache and rule order affect what each layer can see.
Page evidence is normalized
Trakkr stores the path, response result, crawler class, source and available verification evidence.
Citation is observed separately
Tracked prompts and answer samples identify cited URLs. No citation result is inferred from a crawl log.
Referral is observed separately
Analytics records attributable visits from AI services. The exact conversation may remain unknown.
Association is not attribution. A crawl before a citation is a useful investigation path, but timing alone cannot prove the request caused model use, retrieval, ranking, an answer or a human click.
Product views built around decisions
Pages
Compare request volume, crawler classes, citations, referrals, response health and latest activity by normalized page URL.
Map
See which parts of a site receive observed requests and where important pages have no matching activity in the selected window.
Live
Inspect recent classified requests with method, status, path, time, source and collection method.
Actions
Turn access gaps, errors and crawl-without-outcome patterns into review work, without treating them as automatic fixes.
Access
Audit robots.txt, llms.txt, status codes and request outcomes by bot token or class.
Sources
Connect, verify and monitor Cloudflare, Vercel, Netlify, WordPress and other edge or server sources.
| Page | Search fetches | Training fetches | Citations | AI referrals | Next check |
|---|---|---|---|---|---|
| /research/report | Observed | Observed | Observed | Observed | Compare prompt and page evidence |
| /product | No match in window | Observed | No match | No match | Check search access and coverage |
| /docs/old | Blocked response | No match | Historical | No match | Review redirect and WAF order |
Install where the request is visible
| Platform | Trakkr collection method | Availability | What can be missing | Setup |
|---|---|---|---|---|
| Cloudflare | Read-only API token and server-side analytics | All Cloudflare plans can expose AI Crawl Control analytics; detail and retention can vary by plan. | Analytics can be aggregated or sampled. WAF rules can stop requests before Crawl Control records them. | Read steps |
| Vercel | OAuth connection and project-scoped Log Drain | Trakkr log delivery requires a plan that supports the needed Log Drain configuration. | Vercel AI bot rules are a separate control surface. A reverse proxy can reduce Vercel bot-verification context. | Read steps |
| Netlify | OAuth connection and generated Edge Function | The function records matching edge requests and sends the bounded event to Trakkr. | The edge function can only record requests that reach it. Provider or upstream WAF blocks happen first. | Read steps |
| WordPress | Trakkr plugin at the origin | Useful for WordPress sites where an origin plugin is the practical install point. | A CDN, Wordfence or another WAF can stop a request before WordPress sees it. | Read steps |
| CDN and WAF | CloudFront, Akamai, Fastly or generic edge webhook | Choose the earliest safe request layer that exposes the fields your team needs. | Cache, log sampling, rule order and provider retention change what is observable. | Read steps |
| Self-hosted | Next.js, Node or Express, Nginx or OpenResty | Install server or middleware code in your own request path. | Origin logs miss requests served or blocked upstream. Keep secrets out of browser code. | Read steps |
1. Synthetic dashboard test
Confirms Trakkr can store and show a known test event. It does not prove the provider source is installed.
2. Source connection check
Confirms source-specific configuration or delivery health where the integration supports it.
3. Real delivery check
A real, matching request recorded through the installed source is the strongest proof that the live path is working.
Access, identity and status
Policy is not delivery
robots.txt expresses crawler policy. WAF and CDN rules can block first. llms.txt is advisory. A 200 status shows a response was served, not that its content was parsed or used.
Identity needs evidence
Keep the raw user-agent signal, then add official IP lists, reverse DNS or provider verification when available. A copied user-agent is not identity proof.
Status needs context
Record method, response code, collection layer, cache state where available and timestamps. Sampling, retries and time-zone boundaries can explain mismatched totals.
Published crawler research
575,788
classified requests
84
connected sites
314,501
unique URLs
1 Feb 2026
analysis date
Window: 2025-06-11 to 2026-02-01. This panel describes matching request labels in Trakkr's connected-site sample. It is not global crawler market share, verified operator traffic or evidence that a request caused a citation. Dominant e-commerce brand excluded to ensure generalizable patterns
Tools, evidence and next steps
Questions teams ask before installing
AI crawler monitoring records and classifies web requests that match known AI crawler or fetcher signatures. Strong monitoring keeps the request evidence, purpose class, source, path, response result and identity limits visible. It does not assume that every claimed user-agent came from the named operator.
No. A crawl can show that a matching request reached a URL. Citation monitoring is a separate observation. Trakkr compares them by normalized page and time window so teams can investigate patterns, but it does not claim that one caused the other.
No. User-agents are easy to spoof. Use official IP ranges, reverse DNS or provider signatures where available, and keep the verification method with the event. Network verification is stronger evidence, not a guarantee about what happened after the request.
No. OpenAI documents GPTBot and OAI-SearchBot as separate controls. A site can set different policies for possible training use and search discovery. ChatGPT-User is a separate user-triggered fetcher.
No. llms.txt is an advisory file proposed for AI systems. It is not a Google ranking factor, an access-control layer or a guarantee that a provider will read or use a page. robots.txt, HTTP status, network controls and actual request logs remain separate evidence.
They can observe different points in the request path. WAF order, caching, sampling, retries, provider retention, time zones, classification rules and reverse proxies can all change totals. Compare like-for-like windows and preserve the collection method.
A synthetic verification request proves that the dashboard can receive and query a known test event. Source connection checks can prove provider configuration or delivery health. Only a real request seen through the installed source proves that the live ingest path recorded that request.
Turn crawler logs into a page-level evidence loop
Connect the request layer, compare citations and AI referrals, then give technical and content teams a clear next check.