The llms.txt effect
We HTTP-scanned 37,894 AI-cited domains from a corpus of 102,857. 5,035 have llms.txt. The two citation tests give different results. Neither establishes what adding the file would cause.
The Landscape
Based on 37,894 most-cited domains in AI responses
The inverse pattern
The most-cited domains in AI don't have llms.txt. Among the top 50 most-cited sites, only 6% have adopted the standard. As you move down the citation rankings, adoption increases in several broader tiers. This pattern does not explain why sites adopt the file or establish its effect on citations.
The Verdict
Published citation medians: 3 with the file, 3 without
Nearly identical averages
The published snapshot reports averages of 6.8 citations with llms.txt and 6.7 without it. The top 5,000 most-cited domains return p=0.85 in the Mann-Whitney U test. This subset result is not statistically significant at the 0.05 level.
The full 37,894-domain sample returns p<0.001, with reported effect size r=-0.065. That test is significant, while the reported effect is small in magnitude. It must not be described as the same null result.
Who's Adopting
Adoption differs by sector
llms.txt adoption is led by SaaS and developer tools at 24.1%. Government and academic sites sit at just 1.5%, while review sites and reference wikis are at 0%.
The uneven adoption rates leave room for selection bias. The study does not isolate sector, site size, content quality, or existing citation exposure from the decision to publish llms.txt.
The Leaderboard
Most-cited domains that have adopted llms.txt
Most-cited domains that have not adopted llms.txt
Brand Visibility
Composite AI visibility score (0-100)
Mean across all brands analyzed
Based on 205 brands with both audit data and visibility reports
Similar reported visibility scores
Using Trakkr's multi-dimensional visibility scoring - which combines presence, rank, mentions, and sentiment across multiple AI models - brands with llms.txt score 23.15 median visibility versus 23.55 without.
The medians are close. This descriptive comparison does not include confidence intervals or an adjusted causal estimate, so it cannot establish that the difference is noise or that llms.txt has no effect on recommendations.
What This Means
What this actually means
Read these results as a dated observational snapshot, with separate tests for the full sample and its most-cited subset.
Separate the two cohorts
The top-5,000 test is not significant. The full-sample test is significant with a small reported effect size. Reporting either result alone loses useful context.
Association does not establish cause
Sites choose whether to publish llms.txt. Sector, site size, content, and existing citation exposure may differ between groups. This study cannot isolate the effect of adding the file.
The scan does not measure model behavior
Finding a file on a domain does not tell us whether an AI system fetched or used it. Citation counts alone cannot distinguish training, search, or retrieval mechanisms.
Measure the workflow you care about
If you publish llms.txt for a documentation or retrieval workflow, test that workflow directly. A controlled or longitudinal study would be needed to estimate changes in citation outcomes.
Methodology
Published snapshot: 882 citation snapshots, 337,362 citations, and 102,857 unique domains
Ranked domains by AI citation appearances and scanned 37,894 domains with at least two appearances
Async HTTP checks against /llms.txt with content validation to reject HTML error pages and soft 404s
Separate Mann-Whitney U comparisons: top 5,000 cited domains (p=0.85) and full scanned sample (p<0.001, reported r=-0.065)
Separate descriptive comparison of 205 brands with website audit data and visibility reports; not an adjusted causal analysis
Non-parametric test: We used Mann-Whitney U rather than a t-test because citation distributions are heavily right-skewed.
Content validation: HTTP 200 responses were validated to exclude HTML error pages, soft 404s, and login redirects that return 200 status.
Snapshot and scope: The page, charts, and schema use the public snapshot scanned on 2026-03-14. This is a selected sample of AI-cited domains, not an estimate of adoption across the whole web.
Test cohorts: p=0.85 applies to the top 5,000 cited domains. The full 37,894-domain test is reported as p<0.001, with r=-0.065. The JSON stores the full-sample p-value rounded to zero; it is not an exact probability of zero.
Limits: The public release contains aggregate summaries, not domain-level observations, confidence intervals, or an adjusted causal estimate. It does not document the cohort behind the displayed citation means and medians well enough to recompute either test. Do not infer a causal benefit or an exact null effect.
File quality: Among adopters, 89% include a title, 98% contain URLs, and 79% score 4/4 on the published rubric. These describe the files, not their effect on citations.
Continue with this study
See how your brand performs in AI search