F1, precision and recall are estimates from a finite test set, so they need confidence intervals — and F1 has no simple formula for one, so we bootstrap it. Enter the confusion matrix (true/false positives and negatives). Example: TP 80, FP 20, FN 30, TN 370 gives F1 = 0.762 with a 95% interval of 0.695–0.822, precision 0.800 (0.720–0.878) and recall 0.727 (0.644–0.809). The same rates on a test set ten times smaller widen F1’s interval to 0.50–0.93 — small eval sets cannot rank models whose F1 differs by a few points.

Why bootstrap?

Accuracy, precision and recall are proportions and have textbook intervals, but F1 — the harmonic mean of precision and recall, F1 = 2TP / (2TP + FP + FN) — does not have a convenient one. The bootstrap sidesteps the algebra: resample the test items with replacement thousands of times, recompute every metric each time, and read the interval from the middle 95% of those values. Because a confusion matrix records how many items fell in each cell, resampling items is equivalent to resampling the four counts, so you only need the counts.

Worked example

With TP = 80, FP = 20, FN = 30 and TN = 370 (500 items), precision is 0.800, recall 0.727 and F1 0.762. The 95% intervals are 0.720–0.878, 0.644–0.809 and 0.695–0.822. Note that the interval for F1 is about ±0.06 even with 500 items, because only 130 of them are positives or predicted positives — F1 ignores true negatives entirely.

Comparing two models

Non-overlapping intervals suggest a real difference, but overlapping ones do not prove there is none, because both models were scored on the same items. For a proper paired comparison, use the McNemar test on per-item correctness, or bootstrap the difference in F1 by resampling items for both models together.

Common mistakes

  • Quoting F1 to three decimals with no interval on a test set of a few hundred items.
  • Averaging per-class F1 (macro-F1) and treating classes as independent. Bootstrap whole items so that all classes are resampled together.
  • Resampling predictions instead of items. Resample the test items; the model’s predictions travel with them.

Related

AI eval hubCohen's kappaFleiss' kappaKrippendorff's alphaArena Elo ratingsF1 confidence intervalCompare two modelsEval confidence interval (accuracy)

Frequently asked questions

Bootstrap it: resample the test items with replacement many times, recompute F1 each time, and take the 2.5th and 97.5th percentiles as the 95% interval. This calculator does it from the confusion matrix.

F1 only uses true positives, false positives and false negatives. If positives are rare, the effective sample is small even when the test set is large.

A few thousand is standard for 95% percentile intervals. This tool uses up to 5,000 with a fixed seed so the result is reproducible.

Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.