Skip to content

Trakkr behind a WAF

Crawler tracking behind Sucuri, Cloudflare, Wordfence, AWS WAF or another firewall: which install to pick, client IP headers, and safe bot allowlisting.

8 min read

A firewall or security plugin is the most common reason crawler numbers look lower than they should. To a rule engine, an AI crawler looks a lot like a scraper, so the defences that stop scrapers stop crawlers too.

Three things matter when Trakkr sits behind one: which install layer to pick, how the real client IP survives the proxy hops, and how to allowlist a bot without opening a hole.

Pick your install layer

Pick the deepest layer Trakkr can read directly. A CDN-side integration sees traffic an origin-side install never will.

Your setupBest installWhy
Proxied through CloudflareCloudflareReads Cloudflare's own request analytics. Cache, sampling and WAF rule order still affect coverage.
AWS WAF and CloudFrontCloudFront Lambda@EdgeThe template attaches to Viewer Request, before the cache decision.
Behind Sucuri FirewallOrigin-side: WordPress, Next.js, Node or NginxSucuri does not expose a log feed Trakkr can read, so Trakkr reads at the origin.
Behind Imperva, StackPath or another enterprise WAFOrigin-side matching your stackMost enterprise WAF analytics expose no readable feed.
WordPress with Wordfence or Solid SecurityWordPress plugin plus an allowlistThe plugin works at the origin; the allowlist is what gets bots there.

What each install method sees

Requests served from cache never reach your origin, so origin-side integrations do not record them. The same is true of requests the WAF blocks before your stack runs.

Install methodCache hitsRequests the WAF blocked
Cloudflare analyticsDepends on the dataset and planDepends on WAF rule order
CloudFront Lambda@EdgeYes, it runs on Viewer RequestDepends on where AWS WAF runs
Vercel Log DrainYes, Vercel logs cache hitsPartial, depending on edge config
Netlify Edge FunctionDepends on function orderOnly requests that reach the function
WordPress, Next.js, Node, NginxNoNo

How big the origin-side gap is depends on your cache keys, WAF rules and routing. Compare a known window at both layers before you treat one source as authoritative. No method can promise full coverage when something upstream samples, aggregates or drops the request first.

Real client IP detection

Through a proxy, the TCP connection your origin sees comes from the proxy, not the client. The real address rides in a header, and there is no single standard.

HeaderProxy that sets it
CF-Connecting-IPCloudflare
True-Client-IPCloudflare Enterprise, Akamai
X-Sucuri-ClientIPSucuri Firewall
X-Real-IPNginx and generic reverse proxies
X-Forwarded-ForAlmost everything (a comma-separated chain; the client is first)
REMOTE_ADDRThe connection address, used when no proxy header is set

The Trakkr WordPress plugin walks that list in order and takes the first value that parses as an IP:

Text
CF-Connecting-IP → True-Client-IP → X-Sucuri-ClientIP → X-Real-IP → X-Forwarded-For → REMOTE_ADDR

So a WordPress site behind Sucuri still resolves the real bot IP from X-Sucuri-ClientIP, even though every TCP connection arrives from a Sucuri edge server.

Warning
These headers are spoofable. Trakkr records the IP for analytics only, such as country attribution, and never for an access decision. Do not use them for rate limiting or blocking without trust-proxy logic.

The other origin-side integrations ship with the common headers and sometimes need a line adding:

MethodReads by defaultBehind Sucuri or Cloudflare
Next.js middlewarex-forwarded-for, first entryAdd a CF-Connecting-IP or X-Sucuri-ClientIP lookup
Express middlewarereq.ip, with app.set('trust proxy', true)Prepend the proxy-specific header lookup
Nginx and OpenRestyngx.var.remote_addrConfigure real_ip_header and set_real_ip_from for your proxy first
Lambda@Edgerequest.clientIpNothing to do; CloudFront resolves the client IP

Find out whether a bot is being blocked

Send test ping in the Crawlers header writes three labelled synthetic rows straight into Trakkr's store. It never touches your site, so it cannot tell you anything about your firewall. Use it to prove the dashboard works, then look at the layers below.

Two Access tab findings point at a firewall directly. Traffic dropped fires when a bot's requests fell more than 70% against the previous period with no robots rule to explain it, backed by 403 or 429 responses now or sustained traffic before, which is what a new bot mode, WAF rule or rate limit looks like. Access mismatch fires when robots.txt allows a bot but over half of its requests failed across at least ten attempts. The robots.txt panel on the same tab catches the simpler case: a Disallow: / under a bot that your allowlist would otherwise have let through.

To search your firewall's own logs, these are the labels worth grepping for:

Text
GPTBot ChatGPT-User OAI-SearchBot ClaudeBot Claude-User Claude-SearchBot PerplexityBot Perplexity-User Bytespider CCBot Amazonbot MistralAI-User Meta-ExternalFetcher Google-Agent Applebot
Warning
Do not build a permanent bypass from user-agent text. Any client can copy these strings, and forged crawler traffic is common enough that Trakkr has a rule for one specific forged variant. Prefer the provider's verified-bot signal, published IP ranges, reverse DNS or signed requests. If you use a user-agent rule to diagnose, keep it narrow, time-bound and low privilege.

Sucuri Firewall

Open the Sucuri dashboard, then your site, then Security Events, and search the labels above. Confirm the source with operator-published network evidence where it exists. If policy allows the verified crawler, create the narrowest exception Sucuri supports for that source and path, and keep a broad user-agent regex away from sensitive routes.

Tip
Sucuri hardening often blocks /wp-json/. If it does, allowlist /wp-json/trakkr/* under Settings → Access Control, or the WordPress connection will succeed while no crawler rows ever sync.

Cloudflare

Several layers can challenge a crawler, so handle whichever you run.

  • Bot Fight Mode. Check Security Events to see whether the mode challenged the request. The global toggle affects far more than AI crawlers, so change it only within a bounded diagnostic window your policy permits.
  • Bot Management. Prefer Cloudflare's verified-bot and bot-management fields over raw user-agent text. Coverage varies, so compare its classification against the operator's own verification data.
  • Custom WAF rules. Inspect rule order first. If you add a Skip or Allow exception, require stronger verification than a copied user-agent and exclude authentication, checkout and admin paths.

Wordfence

In Wordfence → Firewall → Blocking → Advanced Blocking, check for User-Agent blocks matching AI crawler strings. Then in Wordfence → Live Traffic, filter by user-agent, for example GPTBot or ClaudeBot; rows tagged Blocked are your answer.

Wordfence also disables WordPress Application Passwords by default, which breaks the connection itself rather than the crawler traffic. That switch is under Wordfence → All Options → Brute Force Protection.

Solid Security (formerly iThemes)

In Security → Settings → Network Brute Force Protection → API Settings → Banned Hosts and Banned User Agents, remove wildcard matches such as *bot* or *scraper* that catch AI crawlers. Then check the Firewall section for User Agent filters under System Tweaks.

AWS WAF

Open AWS WAF & Shield → Web ACLs and your ACL. If you use a managed rule group such as AWSManagedRulesCommonRuleSet or AWSManagedRulesBotControlRuleSet, switch the suspect rule to Count temporarily to identify the match, where your security policy allows it. Create an exception only after verifying the source, and never as a higher-priority Allow keyed on a user-agent regex.

Imperva, StackPath and other enterprise WAFs

Most can report the rule and classification that handled a request. Search security events by the claimed user-agent and window, verify the source against operator evidence, then, if policy permits, create a narrow exception tied to the strongest identity signal you have, plus path and method. Providers call this Allow, Bypass or Skip.

Probing from outside

If you cannot tell which layer is responsible, request your own site with a crawler user-agent from a network outside the WAF:

Terminal
"text-[#5b5fc7]">curl "text-accent">-A "Mozilla/5.0 (compatible; GPTBot/1.0; +https://openai.com/gptbot)" \ "text-accent">-I https://your-site.com/

Read the result carefully. This shows how your site treats a request carrying that string from your current network. It does not impersonate or verify GPTBot, and a datacenter IP claiming to be a crawler is exactly the profile a WAF exists to stop. A 403, 503 or challenge page names a policy path worth investigating; a 200 does not prove the real operator can reach the page. The response headers usually reveal which layer answered.

Common questions

The test ping works but real crawlers never appear. What now?

The test ping only proves storage and display; it never leaves Trakkr. Check the connection health message on your source, whether that source sees the production hostname, the robots.txt rules for the missing bot, and then your WAF, bot-management and rate-limit logs in that order.

Should I allowlist AI crawlers by user-agent?

Only as a temporary diagnosis, and never on sensitive paths. A copied user-agent is not identity. Several operators publish IP ranges or signed-request schemes; use those for anything permanent, and refresh them, because the ranges change.

Why do the origin numbers look lower than the CDN numbers?

Because they are counting different things. Cache hits and blocked requests never reach the origin, so an origin-side install cannot see them. That gap is the size of your cache hit rate plus your block rate, not an error.

Does moving my site behind Cloudflare change what Trakkr sees?

Yes, and usually for the better. Once traffic is proxied you can connect the Cloudflare source and see requests your origin never handled. That is why the hosted CMS platforms in the install guide all route through Cloudflare.

Can Trakkr unblock a crawler for me?

No. Trakkr diagnoses the evidence and tells you which layer is likely responsible. Enforcement stays in your robots file, host, CDN or WAF, using that provider's current documentation.

Press ? for keyboard shortcuts