Cross-model consensus benchmark | Trakkr Research
Top-line agreement metrics across the full 8-model comparison set.
Methodology: Built from 797,644 valid comparisons across 44,088 reports and 8 models, covering 6,439,133 model responses in the observed window.
Summary
The 8-model benchmark shows partial overlap, low perfect consensus, and a persistent divergence tail.
Benchmark rows
| Metric | Value | Context |
|---|---|---|
| Average agreement | 43.3% | Mean cross-model agreement rate. |
| Perfect agreement | 4.0% | Only a small share of prompts produce unanimous outcomes. |
| High divergence rate | 14.6% | Prompts in the 0-25% agreement bucket. |
| Average top 3 overlap | 2.8 | Average overlap among top-three results across models. |
Ranked view
| Item | Value | Detail |
|---|---|---|
| Average agreement | 43.3% | The central tendency of cross-model overlap. |
| High-divergence prompts | 14.6% | The share of prompts with the weakest overlap. |
| Perfect agreement | 4.0% | The rare case where all models converge fully. |
| Average top-three overlap | 2.8 | Overlap exists, but not enough to treat lists as identical. |
Related pages
Continue through the same study cluster.
- how often is there perfect consensus across models - Related answer page
- how much do models disagree on brand recommendations - Related answer page
- only four percent of prompts produce perfect consensus - Related fact page
Data & Sources
- Same Question, Different AI, Different Answers - Flagship study behind this page
- Page JSON - Machine-readable companion file