Why a paired test
When two models answer the same items, their scores are not independent: an easy item tends to be right for both, a broken item wrong for both. Those shared outcomes say nothing about which model is better. McNemar’s test throws them away and looks only at the discordant items — where one model is right and the other wrong. If the models were equally good, those disagreements would split roughly 50/50.
How to read the result
The exact p-value is a two-sided binomial test of the discordant split against 50/50. In the example, 52 vs 31 disagreements gives p = 0.028, and the 95% interval for the accuracy gap (0.6 to 7.8 points) excludes zero. Change the split to 30 vs 18 and the gap shrinks to 2.4 points with p = 0.11 — not enough evidence, even though A still scores higher.
Getting the four counts
Score both models on the same items, then count four cells: both right, only A right, only B right, both wrong. You can paste per-item results (one line per item, A then B) and the tool fills the counts for you. Only “only A right” and “only B right” affect the test; the other two set the accuracies.
Common mistakes
- Using a two-proportion z-test. It assumes independent samples and is usually far less powerful here.
- Comparing overlapping confidence intervals. Overlap does not imply no difference for paired data.
- Testing many model variants and reporting the best p-value. With 20 comparisons at α = 0.05, about one will look significant by chance; adjust α (for example Bonferroni) or use a held-out set.
- Mixing sampling settings. If outputs are sampled at temperature > 0, score each model the same way (for example one fixed sample, or majority vote) before counting.
Planning an eval? The eval sample size calculator tells you how many items you need for this test to detect a given gap.
Related
Frequently asked questions
Score both models on the same items and run McNemar's test on the items where they disagree. It is the standard paired test for comparing two classifiers or models on one test set.
The z-test assumes the two accuracies come from independent samples. Two models scored on the same items are paired, so McNemar's test is correct and usually far more powerful.
Use the exact binomial version of McNemar's test, which is valid for any number of disagreements. With very few disagreements, even a large-looking gap may not be significant.
Computed in your browser; formulas verified against SciPy and statsmodels. See the methodology page and the AI & LLM evaluation hub.