Why a paired bootstrap?
When two models are scored on the same items, their errors are linked: hard items trip up both. Comparing two separate confidence intervals ignores this and is far too conservative. The paired bootstrap keeps the link by resampling items and re-scoring both models on every resample, so shared difficulty cancels out. Unlike the McNemar test, which only handles per-item correct/incorrect, it works for any metric computed over a set: F1, macro-F1, precision, BLEU, ROUGE and more.
How it works
- Compute the metric for both models on the full test set; the difference is the observed effect.
- Draw a resample of n items with replacement, and score both models on the same resample.
- Repeat thousands of times. The middle 95% of the differences is the confidence interval; twice the share of differences at or beyond zero (on the far side) is the two-sided p-value.
Our point estimates match scikit-learn’s accuracy_score and f1_score exactly, and the intervals and p-values match an independent NumPy implementation within Monte Carlo error.
Worked example
The preloaded data has 120 items labelled positive, negative or neutral. Model A gets 91.7% right and model B 78.3%, a 13.3-point gap with a 95% interval of 5.0 to 22.5 points (p = 0.004). Switch the metric to macro-F1 and the result is almost identical (+0.134, CI 0.048 to 0.229), because the classes are balanced.
Common mistakes
- Bootstrapping each model separately. That breaks the pairing and inflates the uncertainty.
- Resampling tokens or sentences when the unit is a document. Resample whatever unit is independent, usually a whole test example.
- Running many comparisons and keeping the best. Correct for it with the multiple-comparisons calculator.
- Reporting “p = 0”. With 5,000 resamples the smallest resolvable p-value is 0.0002; report p < 0.0002.
Related
Frequently asked questions
Use a paired bootstrap: resample test items with replacement, score both models on each resample, and use the distribution of differences for a confidence interval and p-value. This works for F1, macro-F1 and other set-level metrics.
McNemar's test only uses per-item correct or incorrect outcomes, so it suits accuracy. The paired bootstrap works for any metric computed over the test set, such as F1 or BLEU.
Several thousand. With 5,000 resamples, the smallest p-value you can report is 1/5000 = 0.0002.
Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.