Study 005

The llms.txt effect

We HTTP-scanned 37,894 AI-cited domains from a corpus of 102,857. 5,035 have llms.txt. The two citation tests give different results. Neither establishes what adding the file would cause.

13.3%
adoption rate
p=0.85
top 5,000 test
37,894
domains scanned
337,362
citations analyzed
Last updated · Mar 14, 2026
[01]

The Landscape

13.3%of AI-cited domains have llms.txt
Adoption by Citation Tier
Top 506.0%
Top 1007.0%
Top 25013.6%
Top 50014.4%
Top 100015.3%
Top 250015.9%
Top 500016.1%
Top 1000015.7%
Top 2500013.7%
Full (37,894)13.3%

Based on 37,894 most-cited domains in AI responses

The inverse pattern

The most-cited domains in AI don't have llms.txt. Among the top 50 most-cited sites, only 6% have adopted the standard. As you move down the citation rankings, adoption increases in several broader tiers. This pattern does not explain why sites adopt the file or establish its effect on citations.

The key question
Many frequently cited domains do not have llms.txt in this scan. The study does not isolate the roles of domain authority, content, retrieval, or training data.
[02]

The Verdict

Average Citations per Domain
With llms.txt
6.8
avg citations
Without llms.txt
6.7
avg citations
Top 5,000, Mann-Whitney U:p=0.85Not Significant
Median (with)
3.0
Median (without)
3.0

Published citation medians: 3 with the file, 3 without

Nearly identical averages

The published snapshot reports averages of 6.8 citations with llms.txt and 6.7 without it. The top 5,000 most-cited domains return p=0.85 in the Mann-Whitney U test. This subset result is not statistically significant at the 0.05 level.

The full 37,894-domain sample returns p<0.001, with reported effect size r=-0.065. That test is significant, while the reported effect is small in magnitude. It must not be described as the same null result.

What the tests establish
These are observational comparisons of different cohorts. They do not show that adding or removing llms.txt causes a change in citations. The subset result does not prove an effect of exactly zero.
[03]

Who's Adopting

24.1%adoption in saas / developer tools
Adoption Rate by Domain Category
SaaS / Developer Tools
97/40324.1%
E-commerce
10/5518.2%
News / Media
52/33215.7%
Social Platforms
84/53615.7%
Government / Academic
9/5811.5%
Reference / Wiki
0/360.0%
Review Sites
0/390.0%

Adoption differs by sector

llms.txt adoption is led by SaaS and developer tools at 24.1%. Government and academic sites sit at just 1.5%, while review sites and reference wikis are at 0%.

The uneven adoption rates leave room for selection bias. The study does not isolate sector, site size, content quality, or existing citation exposure from the decision to publish llms.txt.

Why This Matters
These rates describe the named categories in the scanned sample. They do not establish which site characteristics cause AI citations or measure adoption across the whole web.
[04]

The Leaderboard

With llms.txt

Most-cited domains that have adopted llms.txt

Domain
CitationsBrands
Without llms.txt

Most-cited domains that have not adopted llms.txt

Domain
CitationsBrands
Authority wins
The non-adopter column reads like a who's who of the internet. Reddit, Reuters, Forbes, LinkedIn - these sites dominate AI citations without any llms.txt optimization.
[05]

Brand Visibility

Trakkr Visibility Score Comparison
Median Visibility
With llms.txt
23.1
Without
23.6

Composite AI visibility score (0-100)

Average Visibility
With llms.txt
27.8
Without
26.3

Mean across all brands analyzed

Based on 205 brands with both audit data and visibility reports

Similar reported visibility scores

Using Trakkr's multi-dimensional visibility scoring - which combines presence, rank, mentions, and sentiment across multiple AI models - brands with llms.txt score 23.15 median visibility versus 23.55 without.

The medians are close. This descriptive comparison does not include confidence intervals or an adjusted causal estimate, so it cannot establish that the difference is noise or that llms.txt has no effect on recommendations.

Why This Matters
This analysis cross-references 205 brands that have both website audit data (where we detect llms.txt) and active visibility monitoring. It is a separate brand-level comparison from the domain citation tests.
[06]

What This Means

What this actually means

Read these results as a dated observational snapshot, with separate tests for the full sample and its most-cited subset.

01

Separate the two cohorts

The top-5,000 test is not significant. The full-sample test is significant with a small reported effect size. Reporting either result alone loses useful context.

02

Association does not establish cause

Sites choose whether to publish llms.txt. Sector, site size, content, and existing citation exposure may differ between groups. This study cannot isolate the effect of adding the file.

03

The scan does not measure model behavior

Finding a file on a domain does not tell us whether an AI system fetched or used it. Citation counts alone cannot distinguish training, search, or retrieval mechanisms.

04

Measure the workflow you care about

If you publish llms.txt for a documentation or retrieval workflow, test that workflow directly. A controlled or longitudinal study would be needed to estimate changes in citation outcomes.

Our recommendation
Use the published aggregates to understand adoption and association. Do not treat this snapshot as proof of a citation benefit, proof of no effect, or a test of which AI systems support llms.txt.
[07]

Methodology

Data Pipeline
01
Citation Corpus

Published snapshot: 882 citation snapshots, 337,362 citations, and 102,857 unique domains

02
Domain Ranking

Ranked domains by AI citation appearances and scanned 37,894 domains with at least two appearances

03
llms.txt Detection

Async HTTP checks against /llms.txt with content validation to reject HTML error pages and soft 404s

04
Statistical Testing

Separate Mann-Whitney U comparisons: top 5,000 cited domains (p=0.85) and full scanned sample (p<0.001, reported r=-0.065)

05
Visibility Comparison

Separate descriptive comparison of 205 brands with website audit data and visibility reports; not an adjusted causal analysis

Key Numbers
Domains Scanned
37,894
Brand Snapshots
882
Total Citations
337,362
Top 5,000 Test
p=0.85
Methodology Notes

Non-parametric test: We used Mann-Whitney U rather than a t-test because citation distributions are heavily right-skewed.

Content validation: HTTP 200 responses were validated to exclude HTML error pages, soft 404s, and login redirects that return 200 status.

Snapshot and scope: The page, charts, and schema use the public snapshot scanned on 2026-03-14. This is a selected sample of AI-cited domains, not an estimate of adoption across the whole web.

Test cohorts: p=0.85 applies to the top 5,000 cited domains. The full 37,894-domain test is reported as p<0.001, with r=-0.065. The JSON stores the full-sample p-value rounded to zero; it is not an exact probability of zero.

Limits: The public release contains aggregate summaries, not domain-level observations, confidence intervals, or an adjusted causal estimate. It does not document the cohort behind the displayed citation means and medians well enough to recompute either test. Do not infer a causal benefit or an exact null effect.

File quality: Among adopters, 89% include a title, 98% contain URLs, and 79% score 4/4 on the published rubric. These describe the files, not their effect on citations.

Download public aggregates (CSV) · Source snapshot (JSON)