Current feeds and verified studies · Updated 18h ago

The state of AI search

48.4Mcitation appearances

This report joins current citation, referral, and public-ranking feeds with reviewed studies of crawler behavior, model agreement, page types, query rewrites, and controlled experiments. Each dataset keeps its own sample, observation window, and limits.

Properties
5,037
Models tracked
8
Brands ranked
9,227
Crawler visits
576K
youtube.com1.7M citesen.wikipedia.org887K citesreddit.com842K citeslinkedin.com309K citesgoogle.com302K citesfacebook.com280K citespmc.ncbi.nlm.nih.gov201K citesinstagram.com172K citesyoutube.com1.7M citesen.wikipedia.org887K citesreddit.com842K citeslinkedin.com309K citesgoogle.com302K citesfacebook.com280K citespmc.ncbi.nlm.nih.gov201K citesinstagram.com172K cites

What the data says right now

01
Visible referrals are concentrated

89.6% of detected AI referral pageviews in the current GA4 panel come from ChatGPT. This is panel source mix, not global market share.

02
The models rarely agree

Average pairwise brand-recommendation agreement was 43.3% across 797,644 valid comparisons.

03
Most observed URLs were visited once

88.5% of URLs had exactly one observed crawler visit during the study window. The study does not claim they were never crawled outside that window.

Provenance

The strongest supported findings

This edition separates daily feeds from reviewed snapshots. Every headline below has a definition, sample, date range, source, limit, and reviewer. Verified 2026-08-18.

Verified snapshot2026-08-18

48.4M

Citation appearances

Sum of citation appearance counts in the latest valid snapshot for each tracked brand.

Verified snapshot2026-08-18

90.5%

ChatGPT share of visible AI referrals

ChatGPT-attributed pageviews divided by pageviews attributed to all detected AI referral sources in the latest 30-day window.

Verified snapshot2026-08-18

43.3%

Average cross-model agreement

Average pairwise agreement rate for brand recommendations on matched prompts.

Verified snapshot2026-08-18

88.5%

URLs observed once

Share of URLs with exactly one observed crawler visit during the study window.

ACT IIWho's winning
[01]

Rankings

Who's winning right now?

The public AI 500 projection ranks 9,227 brands from the tracked visibility panel, across 500 industries. It is a Trakkr panel, not a census of all brands.

Unable to load rankings data

ACT IIIWhere attention flows
[02]

Traffic

Who's sending the traffic?

ChatGPT accounts for 89.6% of detected AI referral pageviews in the current GA4 panel. This is visible-referrer source mix, not global AI market share.

7-Day Change
10.0%
30-Day Change
49.3%
Properties
5,037
Panel source mix (30d)
ChatGPT
90%
Claude
6%
Gemini
3%
Perplexity
1%
Grok
0%
DeepSeek
0%

AI referral traffic by source

Rescaled so the highest visible point = 100%

Jun 1Jul 17Sep 1Oct 18Dec 5Jan 23Mar 14May 1Jun 18Sep 20%25%50%75%100%

Aggregated from 5,037 anonymized GA4 properties. Source mix is the share of detected AI referral pageviews over 30 days. It is not global market share.

Explore full AI traffic index
[03]

Citations

Who gets recommended?

youtube.com ranks first with 3.47% of citation appearances. The top three domains account for 7.04%, while the index spans 529,165 domains. This describes the tracked panel and does not establish why a source was selected.

Total Citations
48.4M
Unique Domains
529K
Brands Analyzed
1,643
Look up a domain
Top domain distribution
1
1,680,655
2
886,607
3
842,349
4
309,013
5
301,892
6
280,486
7
201,135
8
171,552
Top Cited Domains
1youtube.com1,680,6553.5%
2en.wikipedia.org886,6071.8%
3reddit.com842,3491.7%
4linkedin.com309,0130.6%
5google.com301,8920.6%
6facebook.com280,4860.6%
7pmc.ncbi.nlm.nih.gov201,1350.4%
8instagram.com171,5520.4%
9forbes.com165,2300.3%
10edmunds.com150,8330.3%
Other
86.8%
Social / UGC
7.2%
Reference
2.0%
Review Sites
1.1%
News & Media
0.7%
Academic
0.7%

You've seen how AI picks its winners. Where does your brand land?

Track your citations, rankings, and movement across ChatGPT, Claude, Gemini, Perplexity, and Google AIOs — updated daily.

Find your brand
ACT IVHow the machines decide
[04]

Model agreement

Do AI models agree?

Average pairwise brand-recommendation agreement was 43.3%. Claude and Deepseek had the highest observed pair score, while Meta and Perplexity had the lowest. The result shows why one model should not be used as a proxy for the full panel.

Average Agreement
43.3%
Perfect Agreement
4.0%
High Divergence
14.6%
Avg Top-3 Overlap
2.8
Model Correlation Matrix
Hover to compare
OpenClauGemiAIOGrokDeepMetaPerp
OpenAI
-
27
21
20
24
25
17
17
Claude
27
-
26
19
35
35
23
15
Gemini
21
26
-
17
24
26
16
12
AIO
20
19
17
-
17
18
12
17
Grok
24
35
24
17
-
31
21
14
Deepseek
25
35
26
18
31
-
22
15
Meta
17
23
16
12
21
22
-
10
Perplexity
17
15
12
17
14
15
10
-
Agreement by Query Type
Comparison
50.4%
How To
45.3%
Alternative
44.1%
Best Of
43.4%
Recommendation
43.1%
[05]

Crawlers

How do they find you?

In this study, OpenAI-associated crawler identities account for 72% of observed visits. Across the URL revisit distribution, 88.5% of URLs had exactly one observed visit during the study window.

Observed visit share by crawler
ChatGPT Training
57.2%
ChatGPT Search
15.1%
ByteSpider
9.2%
Meta
7.9%
Amazon
6.7%
Claude Training
3.8%
Single-Visit Pages
88.5%
of pages visited exactly once
Weekend Reduction
-21%
average weekend crawl reduction
Unique URLs
315K
unique URLs crawled
[06]

Content

What content wins?

In the classified portion of this study, Review / Directory pages have a citation share 4.8 times their observed crawl share, while Homepage pages sit at the other end of the descriptive ratio. Classifier coverage is 58%, and the result is not causal.

Page Type
Crawl ShareCitation Share
Efficiency
Review / Directory
0.3%
1.2%
4.77x
About / Contact
0.4%
1.7%
4.59x
Resource / Report
0.4%
1.4%
3.76x
Service / Use Case
1.4%
2.6%
1.87x
Blog / Editorial
14.4%
20.2%
1.40x
Documentation
0.7%
0.8%
1.13x
Integration
0.4%
0.4%
0.96x
Homepage
10.7%
7.9%
0.74x
[07]

Markdown

Do crawlers prefer Markdown?

Eligible pages are randomized between Markdown and HTML, then compared on observed crawler coverage and live retrieval. Markdown is 76% smaller in the latest reviewed snapshot, but no crawler shows a decisive preference in that snapshot.

Markdown is smaller
76%
median per page
Est. byte savings
33%
across eligible pages
Pages in test
9,033
50/50 split
ChatGPT-User events
2.7M
observed live
Crawler coverage · Markdown vs HTML pages
ChatGPT-User
interaction
MD
82%
HTML
83%
-0.3pp
no signal
OAI-SearchBot
search
MD
86%
HTML
86%
+0.5pp
no signal
ClaudeBot
training
MD
100%
HTML
100%
-0.1pp
no signal
GPTBot
training
MD
5%
HTML
67%
-62.1pp
significant
PerplexityBot
search
MD
19%
HTML
19%
-0.6pp
no signal
[08]

llms.txt

Does llms.txt actually help?

Everyone's rushing to add an llms.txt file. We checked 37,894 AI-cited domains. Only 13.3% have one. The observed citation distributions did not differ significantly (p = 0.85). The honest verdict, for now, is inconclusive.

Adoption Rate
13.3%

of AI-cited domains have adopted llms.txt. While adoption is growing, the file's effect on citation rates remains statistically indistinguishable from noise.

Statistical Test
p = 0.85
Mann-Whitney U - not significant
Domains Scanned
37,894
checked for llms.txt presence
Verdict
Inconclusive
Observational comparison, not a causal test. The question remains open.
ACT VInside the question
[09]

Translation

What happens to your search?

The prompt a person types is almost never the query AI runs. Only 0.17% matched exactly, while 33.0% met the study’s complete-rewrite threshold. Year terms were added in 25.7% of captured pairs.

Complete Rewrite
33.0%
rewritten beyond recognition
Year Injection
25.7%
add current year to queries
Exact Match
0.17%
pass through unchanged
The words AI adds to your search
best
25.3%
list
24.7%
2025
22.6%
top
16.3%
companies
14.8%
brands
7.8%
platforms
6.3%
vendors
5.9%

Share of rewritten queries that gained each word. This is a frequency count, not evidence of a model’s motive.

Illustrative transformations below explain the categories. They are not rows from the withheld prompt dataset.
Year InjectionAdds current year even when you don't ask
IN
best project management tools
OUT
best project management tools 2026
Brand HallucinationInjects brand names into generic queries
IN
best running shoes for flat feet
OUT
best running shoes for flat feet Nike Brooks ASICS
Intent EscalationExpands simple lookups into comparisons
IN
CRM software
OUT
best CRM software for small business comparison review
Audience FabricationInvents demographic and budget constraints
IN
best laptops for students
OUT
best laptops for college students budget under $1000 2026
ACT VIEvidence and limits

What this report can answer

Citation behaviorAvailable

Domain-level appearances, concentration, partial source categories, prompt intent, and observed snapshot history.

Visible AI referralsAvailable

Normalized daily movement, source mix, property coverage, and broad industry change from a qualifying GA4 panel.

Model agreementAvailable

A reviewed point-in-time benchmark of matched brand recommendations across eight tracked models.

Crawler behaviorAvailable

Reviewed aggregate request, revisit, entry, depth, and crawler-identity observations.

Page types and experimentsLimited

Reviewed snapshots with stated classifier coverage, assignment rules, exclusions, and observation windows.

Cross-dataset joinsNot available

The public assets do not support a row-level citation-to-click, crawler-to-citation, or traffic-to-conversion join.

Global market shareNot available

Traffic source mix describes Trakkr’s qualifying GA4 panel, not global AI usage or all website traffic.

Geography and languageNot available

Most public aggregates do not retain a publishable geography or language dimension.

Forecasts and causalityNot available

The report describes observed panels. It does not forecast adoption or claim that a source or page trait caused visibility.

Methodology

Dataset-by-dataset methodology

AI Citation Index

Verified snapshot

Verified 2026-08-18

Records
48,422,016 citation appearances, 2,602,557 unique URLs, 529,165 domains, and 1,643 tracked brands
Dates
2025-10-03 to 2026-08-18
Sampling
Latest valid citation snapshot per tracked brand. Counts are appearance counts within those snapshots, not a census of all AI answers.
Model scope
The public aggregate does not retain provider-level attribution, so model comparisons and response-level citation rates are unavailable.
Geography
Prompt and response geography is not retained in the public aggregate.
Privacy
Downloads include aggregate domain rows seen across at least 3 contributing brands. Raw URLs and customer data are excluded.
Not included
Raw prompts, Raw responses, Customer identities, Account identifiers, Model-level denominators, Geography

AI Search Traffic Index

Verified snapshot

Verified 2026-08-18

Records
4,928 qualifying GA4 properties and 443 daily observations
Dates
2025-06-01 to 2026-08-17
Sampling
A changing panel of connected GA4 properties. Source rows require at least 3 properties and source-by-industry rows require at least 25.
Model scope
Traffic sources are detected from GA4 sessionSource values. This measures visible referrals, not model usage or total AI search volume.
Geography
Geography and device are not published.
Privacy
Public downloads contain normalized trends and aggregate source shares only. Property identifiers and absolute panel volumes are excluded.
Not included
Property identities, Absolute public pageview totals, Landing pages, Conversions, Geography, Device, Citation-to-click joins

Public AI brand rankings

Verified snapshot

Verified 2026-08-18

Records
9,257 ranked brands across 500 industries
Dates
2026-08-17
Sampling
Brands and industries included in the current public rankings projection.
Model scope
Aggregate multi-model visibility ranking. Row-level model evidence is not part of the status response.
Geography
The public status response does not publish geographic coverage.
Privacy
Only public ranking aggregates and approved leaderboard rows are exposed.
Not included
Private brand records, Prompt text, Raw model responses, Customer identities

Model divergence benchmark

Verified snapshot

Verified 2026-08-18

Records
797,644 valid comparisons across 8 models
Dates
2025-08-12 to 2026-03-11
Sampling
Pairwise comparison of brand recommendations for matched prompts with sufficient model coverage.
Model scope
8 tracked models. The public file withholds prompt examples and identifiable rows.
Geography
Not retained in the public aggregate.
Privacy
Only aggregate distributions and rates are published.
Not included
Raw prompts, Raw responses, Brand-level rows, Geography

AI crawler behavior benchmark

Verified snapshot

Verified 2026-08-18

Records
575,788 crawler visits across 84 brands and 314,501 URLs
Dates
2025-06-11 to 2026-02-01
Sampling
Dominant e-commerce brand excluded to ensure generalizable patterns
Model scope
Bot user agents and crawler identities, not generated responses.
Geography
Crawler request geography is not part of this public study.
Privacy
Only aggregate behavior distributions are published.
Not included
Brand identities, Raw request logs, IP addresses, Customer URLs

Page type performance study

Verified snapshot

Verified 2026-08-18

Records
337,362 citations and 11,406,191 crawler visits
Dates
2025-01-01 to 2026-03-14
Sampling
Page classifier coverage is 58%. Unclassified pages remain in the total.
Model scope
Cross-model aggregate only.
Geography
Not retained in the public aggregate.
Privacy
Only aggregate page-type cells are published.
Not included
Raw URLs, Brand identities, Model splits, Unclassified page labels

Markdown crawler experiment

Verified snapshot

Verified 2026-08-18

Records
9,033 eligible pages: 4,517 HTML and 4,516 Markdown
Dates
2026-04-18 to 2026-08-05
Sampling
Canonical paths randomized at page level into HTML and Markdown variants.
Model scope
Observed crawler and live-retrieval cohorts defined in the public experiment file.
Geography
Not published.
Privacy
Only aggregate experiment assignments, results, limitations, and weekly snapshots are published.
Not included
Visitor identities, Raw request logs, IP addresses, Customer data

llms.txt adoption study

Verified snapshot

Verified 2026-08-18

Records
37,894 domains scanned
Dates
2026-03-14
Sampling
Domains observed in the Trakkr citation corpus and scanned for an accessible llms.txt file.
Model scope
Citation aggregate does not provide a publishable model split.
Geography
Global public domains; country mix is not published.
Privacy
Only aggregate adoption and outcome distributions are published.
Not included
Named adopters, Named non-adopters, Customer identities

Query translation study

Verified snapshot

Verified 2026-08-18

Records
11,521 prompt-query pairs
Dates
2026-01-22
Sampling
Prompt-query pairs with both original prompt and generated retrieval query available for comparison.
Model scope
Aggregate retrieval transformations. The public file withholds example text.
Geography
Not retained in the public aggregate.
Privacy
Examples are redacted and withheld; only aggregate transformation rates are published.
Not included
Raw prompts, Raw queries, Named brands, Language examples

Flagship metric ledger

Open any metric to inspect its calculation, sample, scope, and limits.

Definition
Sum of citation appearance counts in the latest valid snapshot for each tracked brand.
Numerator
Not applicable
Denominator
Not applicable
Sample
1,643 tracked brands
Dates
2025-10-03 to 2026-08-18
Model scope
The public aggregate does not retain provider-level attribution, so model comparisons and response-level citation rates are unavailable.
Geography
Prompt and response geography is not retained in the public aggregate.
Limits and confidence
Not unique citations and not a census of all AI answers. A URL can appear more than once. Direct aggregate count from the verified public snapshot.
Review
Verified 2026-08-18. Reviewed by Mack Grenfell.

Download the aggregate data

Aggregate, privacy-filtered files. No customer identities, prompts, account IDs, raw URLs, or absolute traffic volumes.

Definitions

Citation appearance
One observed instance of a source URL appearing in a tracked AI response snapshot. It is not necessarily a unique URL or a unique answer.
Unique cited domain
A normalized hostname with at least one citation appearance in the current brand snapshots.
Visible AI referral
A GA4 sessionSource value matching a known AI assistant referrer. Visits without a visible referrer cannot be attributed.
Traffic index
A smoothed series normalized to a peak value. It shows movement, not the public panel’s absolute traffic volume.
Source mix
A source’s share of detected AI-attributed pageviews in the selected panel and period. It is not global AI market share.
Verified snapshot
A point-in-time value regenerated from an attributable public aggregate and checked against its numerator, denominator, and coverage metadata.

Update history

2026-08-18

Flagship evidence refresh

Rebuilt the three public research surfaces from 48,422,016 citation appearances and 4,928 qualifying GA4 properties.

2026-05-06

Markdown experiment added

Added the randomized 9,033-page crawler experiment as a reviewed State of AI Search dataset.

2026-03-30

Longitudinal citation study added

Added citation persistence and volatility findings with an explicit 177-day citation observation window.

Corrections

2026-08-18

Citation headline and Wikipedia share

Removed the stale 1.3M citation and 17% Wikipedia claims. The page now resolves current values from the verified Citation Index and records the snapshot date.

2026-08-18

Traffic loading values

Removed invented placeholder growth and source-share values. Loading and failed requests now render as unavailable.

2026-08-18

Live-data description

Replaced “every number is live” with dataset-level status because several reviewed studies are point-in-time snapshots.

Written and reviewed by Mack Grenfell. Verified 2026-08-18. The evidence manifest and downloads are rebuilt with npm run research:build, then checked with provenance and privacy validators.

How we measure this

This report combines current aggregate feeds with reviewed point-in-time studies. The citation feed covers 48.4M appearances and the referral panel covers 5,037 qualifying GA4 properties. The crawler, model, page-type, llms.txt, query-translation, and experiment sections keep their original study windows. Each material figure is marked as a live feed or verified snapshot in the evidence ledger above.

Current feeds
Citation, traffic, and ranking feeds refresh independently
Reviewed studies
Point-in-time datasets retain their own observation window
Open
Aggregate evidence and downloads are public
Bounded
Unavailable joins, geography, and causal claims stay unavailable

You've seen the whole map. Now find your brand on it.

Track how you appear across ChatGPT, Claude, Gemini, Perplexity, and Google AIOs — your citations, rankings, and movement, updated daily.

Start free

14-day free trial · Cancel anytime