Are general and best-of prompts more volatile than comparisons? | Trakkr Research
Yes. Comparison prompts average 50.4% agreement, while general prompts average 42.2% and best-of prompts carry a 14.8% high-divergence rate.
Methodology: Built from 797,644 valid comparisons across 44,088 reports and 8 models, covering 6,439,133 model responses in the observed window.
Direct Answer
Yes. Comparison prompts average 50.4% agreement, while general prompts average 42.2% and best-of prompts carry a 14.8% high-divergence rate.
What this means
Operators must allocate resources differently based on query type, as high-divergence categories require multi-model optimization rather than single-platform focus.
Evidence table
| Metric | Value | Why it matters |
|---|---|---|
| Comparison-query agreement | 50.4% | Comparison prompts produce the highest average agreement. |
| General-query agreement | 42.2% | General prompts are less stable across models. |
| Best-of high divergence | 14.8% | Best-of prompts frequently split models. |
Frequently Asked Questions
Which prompt type produces the highest agreement across models?
Comparison prompts produce the highest average agreement at 50.4%.
How often do best-of prompts cause models to split?
Best-of prompts carry a 14.8% high-divergence rate, frequently splitting models.
What to do next
Related pages
Continue through the same study cluster.
- what does an average top three overlap of two point eight mean - Related answer page
- should you use one model as a proxy for all ai visibility - Related answer page
- comparison prompts are the most stable query class - Related fact page
- query class agreement tracker - Related tracker page
Data & Sources
- Same Question, Different AI, Different Answers - Flagship study behind this page
- Page JSON - Machine-readable companion file