The paired bootstrap tests whether one model really beats another on the same test set, for any metric — including F1 and macro-F1, which have no simple formula. It resamples test items with replacement, re-scores both models on each resample, and reads a confidence interval and p-value from the thousands of resulting differences. Example: on 120 sentiment items, model A scores 91.7% accuracy and model B 78.3%. The difference, +13.3 points, has a 95% CI of 5.0 to 22.5 and p = 0.004, so A is genuinely better. This is the standard significance test in NLP papers (Koehn 2004; Berg-Kirkpatrick et al. 2012).

Commas or tabs; labels can be words or numbers. Both models must be scored on the same items, in the same order.

Why a paired bootstrap?

When two models are scored on the same items, their errors are linked: hard items trip up both. Comparing two separate confidence intervals ignores this and is far too conservative. The paired bootstrap keeps the link by resampling items and re-scoring both models on every resample, so shared difficulty cancels out. Unlike the McNemar test, which only handles per-item correct/incorrect, it works for any metric computed over a set: F1, macro-F1, precision, BLEU, ROUGE and more.

How it works

  1. Compute the metric for both models on the full test set; the difference is the observed effect.
  2. Draw a resample of n items with replacement, and score both models on the same resample.
  3. Repeat thousands of times. The middle 95% of the differences is the confidence interval; twice the share of differences at or beyond zero (on the far side) is the two-sided p-value.

Our point estimates match scikit-learn’s accuracy_score and f1_score exactly, and the intervals and p-values match an independent NumPy implementation within Monte Carlo error.

Worked example

The preloaded data has 120 items labelled positive, negative or neutral. Model A gets 91.7% right and model B 78.3%, a 13.3-point gap with a 95% interval of 5.0 to 22.5 points (p = 0.004). Switch the metric to macro-F1 and the result is almost identical (+0.134, CI 0.048 to 0.229), because the classes are balanced.

Common mistakes

  • Bootstrapping each model separately. That breaks the pairing and inflates the uncertainty.
  • Resampling tokens or sentences when the unit is a document. Resample whatever unit is independent, usually a whole test example.
  • Running many comparisons and keeping the best. Correct for it with the multiple-comparisons calculator.
  • Reporting “p = 0”. With 5,000 resamples the smallest resolvable p-value is 0.0002; report p < 0.0002.

Related

AI eval hubCohen's kappaFleiss' kappaKrippendorff's alphaArena Elo ratingsF1 confidence intervalCompare two modelsMultiple-comparisons correction

Frequently asked questions

Use a paired bootstrap: resample test items with replacement, score both models on each resample, and use the distribution of differences for a confidence interval and p-value. This works for F1, macro-F1 and other set-level metrics.

McNemar's test only uses per-item correct or incorrect outcomes, so it suits accuracy. The paired bootstrap works for any metric computed over the test set, such as F1 or BLEU.

Several thousand. With 5,000 resamples, the smallest p-value you can report is 1/5000 = 0.0002.

Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.