Are general and best-of prompts more volatile than comparisons? | Trakkr Research

Yes. Comparison prompts average 50.4% agreement, while general prompts average 42.2% and best-of prompts carry a 14.8% high-divergence rate.

Methodology: Built from 797,644 valid comparisons across 44,088 reports and 8 models, covering 6,439,133 model responses in the observed window.

Direct Answer

Yes. Comparison prompts average 50.4% agreement, while general prompts average 42.2% and best-of prompts carry a 14.8% high-divergence rate.

What this means

Operators must allocate resources differently based on query type, as high-divergence categories require multi-model optimization rather than single-platform focus.

Evidence table

Metric Value Why it matters
Comparison-query agreement 50.4% Comparison prompts produce the highest average agreement.
General-query agreement 42.2% General prompts are less stable across models.
Best-of high divergence 14.8% Best-of prompts frequently split models.

Frequently Asked Questions

Which prompt type produces the highest agreement across models?

Comparison prompts produce the highest average agreement at 50.4%.

How often do best-of prompts cause models to split?

Best-of prompts carry a 14.8% high-divergence rate, frequently splitting models.

What to do next

Related pages

Continue through the same study cluster.

Data & Sources