Cross-model consensus benchmark | Trakkr Research

Top-line agreement metrics across the full 8-model comparison set.

Methodology: Built from 797,644 valid comparisons across 44,088 reports and 8 models, covering 6,439,133 model responses in the observed window.

Summary

The 8-model benchmark shows partial overlap, low perfect consensus, and a persistent divergence tail.

Benchmark rows

Metric Value Context
Average agreement 43.3% Mean cross-model agreement rate.
Perfect agreement 4.0% Only a small share of prompts produce unanimous outcomes.
High divergence rate 14.6% Prompts in the 0-25% agreement bucket.
Average top 3 overlap 2.8 Average overlap among top-three results across models.

Ranked view

Item Value Detail
Average agreement 43.3% The central tendency of cross-model overlap.
High-divergence prompts 14.6% The share of prompts with the weakest overlap.
Perfect agreement 4.0% The rare case where all models converge fully.
Average top-three overlap 2.8 Overlap exists, but not enough to treat lists as identical.

Related pages

Continue through the same study cluster.

Data & Sources