Why bootstrap?
Accuracy, precision and recall are proportions and have textbook intervals, but F1 — the harmonic mean of precision and recall, F1 = 2TP / (2TP + FP + FN) — does not have a convenient one. The bootstrap sidesteps the algebra: resample the test items with replacement thousands of times, recompute every metric each time, and read the interval from the middle 95% of those values. Because a confusion matrix records how many items fell in each cell, resampling items is equivalent to resampling the four counts, so you only need the counts.
Worked example
With TP = 80, FP = 20, FN = 30 and TN = 370 (500 items), precision is 0.800, recall 0.727 and F1 0.762. The 95% intervals are 0.720–0.878, 0.644–0.809 and 0.695–0.822. Note that the interval for F1 is about ±0.06 even with 500 items, because only 130 of them are positives or predicted positives — F1 ignores true negatives entirely.
Comparing two models
Non-overlapping intervals suggest a real difference, but overlapping ones do not prove there is none, because both models were scored on the same items. For a proper paired comparison, use the McNemar test on per-item correctness, or bootstrap the difference in F1 by resampling items for both models together.
Common mistakes
- Quoting F1 to three decimals with no interval on a test set of a few hundred items.
- Averaging per-class F1 (macro-F1) and treating classes as independent. Bootstrap whole items so that all classes are resampled together.
- Resampling predictions instead of items. Resample the test items; the model’s predictions travel with them.
Related
Frequently asked questions
Bootstrap it: resample the test items with replacement many times, recompute F1 each time, and take the 2.5th and 97.5th percentiles as the 95% interval. This calculator does it from the confusion matrix.
F1 only uses true positives, false positives and false negatives. If positives are rare, the effective sample is small even when the test set is large.
A few thousand is standard for 95% percentile intervals. This tool uses up to 5,000 with a fixed seed so the result is reproducible.
Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.