AI crawler monitoring

See the request. Check the citation. Measure the visit.

Trakkr records AI crawler and fetcher activity by page, then compares it with separately observed citations, AI referral traffic, access signals and next actions. You get a useful funnel without pretending a crawl caused an answer.

Page evidence
Crawler requestObservedOAI-SearchBot signature
Response200HTML served at edge
CitationObserved laterSeparate answer sample
AI referral3 sessionsAttributed by referrer or campaign
ConclusionReviewAssociation, not causation

Crawler monitoring evidence model

What is recorded?

A matching HTTP request, normalized page, time, method, response status, referrer when present, crawler class, user-agent signal, source and collection method.

What is verified?

The UI distinguishes a signature match from stronger network or provider evidence. User-agent alone never proves who sent a request.

What is joined?

Crawler events, citations, AI referrals and optional search data are compared on normalized page URLs and aligned windows, with their source and limits intact.

[01]

Separate purpose before policy

Training crawlers

GPTBot, ClaudeBot

Fetch pages that may be used to improve models. Their robots controls are separate from search crawler controls where the operator documents separate tokens.

Search and indexing crawlers

OAI-SearchBot, Claude-SearchBot, PerplexityBot

Build or refresh retrieval indexes used by AI search products. A successful request can support access or eligibility, but does not predict ranking or citation.

User-requested fetchers

ChatGPT-User, Claude-User, Perplexity-User

Retrieve a page in response to a user action. A request does not reveal the conversation, prove the page was read, or prove that an answer cited it.

Control tokens

Google-Extended

Set a robots.txt policy but may not appear as a distinct HTTP user-agent. Google says Google-Extended does not affect Search inclusion or ranking.

[02]

The crawler to citation to referral funnel

[1]

Request reaches a layer

Edge, CDN, WAF or origin records a matching request. Cache and rule order affect what each layer can see.

[2]

Page evidence is normalized

Trakkr stores the path, response result, crawler class, source and available verification evidence.

[3]

Citation is observed separately

Tracked prompts and answer samples identify cited URLs. No citation result is inferred from a crawl log.

[4]

Referral is observed separately

Analytics records attributable visits from AI services. The exact conversation may remain unknown.

Association is not attribution. A crawl before a citation is a useful investigation path, but timing alone cannot prove the request caused model use, retrieval, ranking, an answer or a human click.

[03]

Product views built around decisions

Pages

Compare request volume, crawler classes, citations, referrals, response health and latest activity by normalized page URL.

Map

See which parts of a site receive observed requests and where important pages have no matching activity in the selected window.

Live

Inspect recent classified requests with method, status, path, time, source and collection method.

Actions

Turn access gaps, errors and crawl-without-outcome patterns into review work, without treating them as automatic fixes.

Access

Audit robots.txt, llms.txt, status codes and request outcomes by bot token or class.

Sources

Connect, verify and monitor Cloudflare, Vercel, Netlify, WordPress and other edge or server sources.

Illustrative page-level comparison. Values and states shown here explain the interface; they are not customer data.
PageSearch fetchesTraining fetchesCitationsAI referralsNext check
/research/reportObservedObservedObservedObservedCompare prompt and page evidence
/productNo match in windowObservedNo matchNo matchCheck search access and coverage
/docs/oldBlocked responseNo matchHistoricalNo matchReview redirect and WAF order
[04]

Install where the request is visible

Supported crawler monitoring install methods and known limitations
PlatformTrakkr collection methodAvailabilityWhat can be missingSetup
CloudflareRead-only API token and server-side analyticsAll Cloudflare plans can expose AI Crawl Control analytics; detail and retention can vary by plan.Analytics can be aggregated or sampled. WAF rules can stop requests before Crawl Control records them.Read steps
VercelOAuth connection and project-scoped Log DrainTrakkr log delivery requires a plan that supports the needed Log Drain configuration.Vercel AI bot rules are a separate control surface. A reverse proxy can reduce Vercel bot-verification context.Read steps
NetlifyOAuth connection and generated Edge FunctionThe function records matching edge requests and sends the bounded event to Trakkr.The edge function can only record requests that reach it. Provider or upstream WAF blocks happen first.Read steps
WordPressTrakkr plugin at the originUseful for WordPress sites where an origin plugin is the practical install point.A CDN, Wordfence or another WAF can stop a request before WordPress sees it.Read steps
CDN and WAFCloudFront, Akamai, Fastly or generic edge webhookChoose the earliest safe request layer that exposes the fields your team needs.Cache, log sampling, rule order and provider retention change what is observable.Read steps
Self-hostedNext.js, Node or Express, Nginx or OpenRestyInstall server or middleware code in your own request path.Origin logs miss requests served or blocked upstream. Keep secrets out of browser code.Read steps

1. Synthetic dashboard test

Confirms Trakkr can store and show a known test event. It does not prove the provider source is installed.

2. Source connection check

Confirms source-specific configuration or delivery health where the integration supports it.

3. Real delivery check

A real, matching request recorded through the installed source is the strongest proof that the live path is working.

[05]

Access, identity and status

Policy is not delivery

robots.txt expresses crawler policy. WAF and CDN rules can block first. llms.txt is advisory. A 200 status shows a response was served, not that its content was parsed or used.

Identity needs evidence

Keep the raw user-agent signal, then add official IP lists, reverse DNS or provider verification when available. A copied user-agent is not identity proof.

Status needs context

Record method, response code, collection layer, cache state where available and timestamps. Sampling, retries and time-zone boundaries can explain mismatched totals.

Trakkr is a monitoring and diagnosis layer. It does not claim to control or block a provider merely because it can show access findings. Enforcement remains in robots.txt, the CDN, WAF, host or source-specific integration where supported.
[06]

Published crawler research

575,788

classified requests

84

connected sites

314,501

unique URLs

1 Feb 2026

analysis date

Window: 2025-06-11 to 2026-02-01. This panel describes matching request labels in Trakkr's connected-site sample. It is not global crawler market share, verified operator traffic or evidence that a request caused a citation. Dominant e-commerce brand excluded to ensure generalizable patterns

[07]

Tools, evidence and next steps

[08]

Questions teams ask before installing

AI crawler monitoring records and classifies web requests that match known AI crawler or fetcher signatures. Strong monitoring keeps the request evidence, purpose class, source, path, response result and identity limits visible. It does not assume that every claimed user-agent came from the named operator.

Turn crawler logs into a page-level evidence loop

Connect the request layer, compare citations and AI referrals, then give technical and content teams a clear next check.