Why disagreement drives the sample size
A paired comparison learns only from items where the models disagree. If two models get the same items right and wrong most of the time, the few disagreements carry a clean signal and you need relatively few items. If their errors were unrelated, disagreements would be common and noisy, and pairing would give almost no benefit — the paired and unpaired numbers converge (1,097 vs 1,091 in the default example).
How to estimate the disagreement rate
Run both models on a pilot of 100–200 items and count the share where exactly one is correct. Related models (two checkpoints, two prompts on the same base model) often disagree on only 5–15% of items. Leaving the field blank assumes independent errors, which gives a safe upper bound.
The formula
With p10 and p01 the probabilities that only A or only B is correct, ψ = p10 + p01 and δ = p10 − p01 (the accuracy gap), Connor’s (1987) formula for McNemar’s test is
n = (z1−α/2√ψ + z1−β√(ψ − δ²))² / δ²
We checked it by simulation: across 20,000 simulated evals at the computed sizes, the test reached 79.7–80.1% power against an 80% target.
Common mistakes
- Sizing an eval for one model’s margin of error when the real question is a comparison — they give different answers.
- Choosing the gap after seeing results. Decide the smallest gap that would change your decision before running the eval.
- Counting sampled repeats as extra items. Five samples of one prompt are not five independent items.
Related
Frequently asked questions
It depends on the accuracy gap you want to detect and how often the two models disagree. Detecting 80% vs 75% needs about 1,097 items assuming independent errors, but only 469 if the models disagree on 15% of items.
Scoring both models on the same items cancels out item difficulty. Only the items where the models disagree carry information, so when disagreements are rare the comparison is precise with fewer items.
Measure it on a pilot of 100 to 200 items. Leaving it blank assumes independent errors, which is a conservative upper bound.
Computed in your browser; formulas verified against SciPy and statsmodels. See the methodology page and the AI & LLM evaluation hub.