Trakkr behind a WAF
Crawler tracking behind Sucuri, Cloudflare, Wordfence, AWS WAF or another firewall: which install to pick, client IP headers, and safe bot allowlisting.
A firewall or security plugin is the most common reason crawler numbers look lower than they should. To a rule engine, an AI crawler looks a lot like a scraper, so the defences that stop scrapers stop crawlers too.
Three things matter when Trakkr sits behind one: which install layer to pick, how the real client IP survives the proxy hops, and how to allowlist a bot without opening a hole.
Pick your install layer
Pick the deepest layer Trakkr can read directly. A CDN-side integration sees traffic an origin-side install never will.
| Your setup | Best install | Why |
|---|---|---|
| Proxied through Cloudflare | Cloudflare | Reads Cloudflare's own request analytics. Cache, sampling and WAF rule order still affect coverage. |
| AWS WAF and CloudFront | CloudFront Lambda@Edge | The template attaches to Viewer Request, before the cache decision. |
| Behind Sucuri Firewall | Origin-side: WordPress, Next.js, Node or Nginx | Sucuri does not expose a log feed Trakkr can read, so Trakkr reads at the origin. |
| Behind Imperva, StackPath or another enterprise WAF | Origin-side matching your stack | Most enterprise WAF analytics expose no readable feed. |
| WordPress with Wordfence or Solid Security | WordPress plugin plus an allowlist | The plugin works at the origin; the allowlist is what gets bots there. |
What each install method sees
Requests served from cache never reach your origin, so origin-side integrations do not record them. The same is true of requests the WAF blocks before your stack runs.
| Install method | Cache hits | Requests the WAF blocked |
|---|---|---|
| Cloudflare analytics | Depends on the dataset and plan | Depends on WAF rule order |
| CloudFront Lambda@Edge | Yes, it runs on Viewer Request | Depends on where AWS WAF runs |
| Vercel Log Drain | Yes, Vercel logs cache hits | Partial, depending on edge config |
| Netlify Edge Function | Depends on function order | Only requests that reach the function |
| WordPress, Next.js, Node, Nginx | No | No |
How big the origin-side gap is depends on your cache keys, WAF rules and routing. Compare a known window at both layers before you treat one source as authoritative. No method can promise full coverage when something upstream samples, aggregates or drops the request first.
Real client IP detection
Through a proxy, the TCP connection your origin sees comes from the proxy, not the client. The real address rides in a header, and there is no single standard.
| Header | Proxy that sets it |
|---|---|
CF-Connecting-IP | Cloudflare |
True-Client-IP | Cloudflare Enterprise, Akamai |
X-Sucuri-ClientIP | Sucuri Firewall |
X-Real-IP | Nginx and generic reverse proxies |
X-Forwarded-For | Almost everything (a comma-separated chain; the client is first) |
REMOTE_ADDR | The connection address, used when no proxy header is set |
The Trakkr WordPress plugin walks that list in order and takes the first value that parses as an IP:
CF-Connecting-IP → True-Client-IP → X-Sucuri-ClientIP → X-Real-IP → X-Forwarded-For → REMOTE_ADDRSo a WordPress site behind Sucuri still resolves the real bot IP from X-Sucuri-ClientIP, even though every TCP connection arrives from a Sucuri edge server.
The other origin-side integrations ship with the common headers and sometimes need a line adding:
| Method | Reads by default | Behind Sucuri or Cloudflare |
|---|---|---|
| Next.js middleware | x-forwarded-for, first entry | Add a CF-Connecting-IP or X-Sucuri-ClientIP lookup |
| Express middleware | req.ip, with app.set('trust proxy', true) | Prepend the proxy-specific header lookup |
| Nginx and OpenResty | ngx.var.remote_addr | Configure real_ip_header and set_real_ip_from for your proxy first |
| Lambda@Edge | request.clientIp | Nothing to do; CloudFront resolves the client IP |
Find out whether a bot is being blocked
Send test ping in the Crawlers header writes three labelled synthetic rows straight into Trakkr's store. It never touches your site, so it cannot tell you anything about your firewall. Use it to prove the dashboard works, then look at the layers below.
Two Access tab findings point at a firewall directly. Traffic dropped fires when a bot's requests fell more than 70% against the previous period with no robots rule to explain it, backed by 403 or 429 responses now or sustained traffic before, which is what a new bot mode, WAF rule or rate limit looks like. Access mismatch fires when robots.txt allows a bot but over half of its requests failed across at least ten attempts. The robots.txt panel on the same tab catches the simpler case: a Disallow: / under a bot that your allowlist would otherwise have let through.
To search your firewall's own logs, these are the labels worth grepping for:
GPTBot
ChatGPT-User
OAI-SearchBot
ClaudeBot
Claude-User
Claude-SearchBot
PerplexityBot
Perplexity-User
Bytespider
CCBot
Amazonbot
MistralAI-User
Meta-ExternalFetcher
Google-Agent
ApplebotSucuri Firewall
Open the Sucuri dashboard, then your site, then Security Events, and search the labels above. Confirm the source with operator-published network evidence where it exists. If policy allows the verified crawler, create the narrowest exception Sucuri supports for that source and path, and keep a broad user-agent regex away from sensitive routes.
/wp-json/. If it does, allowlist /wp-json/trakkr/* under Settings → Access Control, or the WordPress connection will succeed while no crawler rows ever sync.Cloudflare
Several layers can challenge a crawler, so handle whichever you run.
- Bot Fight Mode. Check Security Events to see whether the mode challenged the request. The global toggle affects far more than AI crawlers, so change it only within a bounded diagnostic window your policy permits.
- Bot Management. Prefer Cloudflare's verified-bot and bot-management fields over raw user-agent text. Coverage varies, so compare its classification against the operator's own verification data.
- Custom WAF rules. Inspect rule order first. If you add a Skip or Allow exception, require stronger verification than a copied user-agent and exclude authentication, checkout and admin paths.
Wordfence
In Wordfence → Firewall → Blocking → Advanced Blocking, check for User-Agent blocks matching AI crawler strings. Then in Wordfence → Live Traffic, filter by user-agent, for example GPTBot or ClaudeBot; rows tagged Blocked are your answer.
Wordfence also disables WordPress Application Passwords by default, which breaks the connection itself rather than the crawler traffic. That switch is under Wordfence → All Options → Brute Force Protection.
Solid Security (formerly iThemes)
In Security → Settings → Network Brute Force Protection → API Settings → Banned Hosts and Banned User Agents, remove wildcard matches such as *bot* or *scraper* that catch AI crawlers. Then check the Firewall section for User Agent filters under System Tweaks.
AWS WAF
Open AWS WAF & Shield → Web ACLs and your ACL. If you use a managed rule group such as AWSManagedRulesCommonRuleSet or AWSManagedRulesBotControlRuleSet, switch the suspect rule to Count temporarily to identify the match, where your security policy allows it. Create an exception only after verifying the source, and never as a higher-priority Allow keyed on a user-agent regex.
Imperva, StackPath and other enterprise WAFs
Most can report the rule and classification that handled a request. Search security events by the claimed user-agent and window, verify the source against operator evidence, then, if policy permits, create a narrow exception tied to the strongest identity signal you have, plus path and method. Providers call this Allow, Bypass or Skip.
Probing from outside
If you cannot tell which layer is responsible, request your own site with a crawler user-agent from a network outside the WAF:
"text-[#5b5fc7]">curl "text-accent">-A "Mozilla/5.0 (compatible; GPTBot/1.0; +https://openai.com/gptbot)" \
"text-accent">-I https://your-site.com/Read the result carefully. This shows how your site treats a request carrying that string from your current network. It does not impersonate or verify GPTBot, and a datacenter IP claiming to be a crawler is exactly the profile a WAF exists to stop. A 403, 503 or challenge page names a policy path worth investigating; a 200 does not prove the real operator can reach the page. The response headers usually reveal which layer answered.
Common questions
The test ping works but real crawlers never appear. What now?
The test ping only proves storage and display; it never leaves Trakkr. Check the connection health message on your source, whether that source sees the production hostname, the robots.txt rules for the missing bot, and then your WAF, bot-management and rate-limit logs in that order.
Should I allowlist AI crawlers by user-agent?
Only as a temporary diagnosis, and never on sensitive paths. A copied user-agent is not identity. Several operators publish IP ranges or signed-request schemes; use those for anything permanent, and refresh them, because the ranges change.
Why do the origin numbers look lower than the CDN numbers?
Because they are counting different things. Cache hits and blocked requests never reach the origin, so an origin-side install cannot see them. That gap is the size of your cache hit rate plus your block rate, not an error.
Does moving my site behind Cloudflare change what Trakkr sees?
Yes, and usually for the better. Once traffic is proxied you can connect the Cloudflare source and see requests your origin never handled. That is why the hosted CMS platforms in the install guide all route through Cloudflare.
Can Trakkr unblock a crawler for me?
No. Trakkr diagnoses the evidence and tells you which layer is likely responsible. Enforcement stays in your robots file, host, CDN or WAF, using that provider's current documentation.
