When you run many tests, some will look significant by chance, so their p-values must be adjusted. Bonferroni and Holm control the family-wise error rate (the chance of any false positive); Benjamini–Hochberg controls the false discovery rate (the expected share of false positives among your discoveries) and keeps more power. Example: 10 prompt variants are each tested against a baseline. Uncorrected, 7 of 10 have p < 0.05. After correction, 1 survives Bonferroni or Holm and 3 survive Benjamini–Hochberg.

Why correct for multiple comparisons

At α = 0.05, each test has a 5% chance of a false positive when nothing is really going on. Run 20 independent tests and the chance of at least one false positive is 1 − 0.9520 ≈ 64%. This happens constantly in practice: comparing many prompts or checkpoints, testing many metrics, analysing many subgroups, or screening many genes.

Which method to use

  • Bonferroni — multiply each p by the number of tests. Simple and strict; controls the family-wise error rate (FWER).
  • Holm — a step-down version of Bonferroni that is never less powerful. Use it instead of Bonferroni whenever you need FWER control.
  • Šidák — slightly less conservative than Bonferroni when tests are independent.
  • Benjamini–Hochberg (BH) — controls the false discovery rate (FDR): among results you call significant, at most α are expected to be false. Much more powerful with many tests; standard in genomics and exploratory analyses.
  • Benjamini–Yekutieli (BY) — FDR control that holds under any dependence between tests, at a cost in power.

Worked example

Ten prompt variants are compared with a baseline, giving p-values from 0.001 to 0.62. Seven are below 0.05. Bonferroni multiplies by 10, leaving only v1 (adjusted p = 0.01). Holm gives the same answer here (v2’s adjusted p is 0.072). Benjamini–Hochberg keeps v1, v2 and v3 (adjusted p = 0.01, 0.04 and 0.04). If the goal is to shortlist promising prompts for a follow-up test, BH is appropriate; if a false claim is costly, use Holm.

Common mistakes

  • Correcting only the tests that looked interesting. Count every test you ran, including the ones you did not report.
  • Using Bonferroni when Holm is available. Holm controls the same error rate with more power.
  • Treating FDR and FWER as the same guarantee. BH allows some false positives among discoveries; it is a different promise.

Related

Compare two models (McNemar)Paired bootstrap testType I & II errorsInterpreting p-valuesAll tools

Frequently asked questions

Use Bonferroni or Holm when any false positive is costly, because they control the chance of even one. Use Benjamini-Hochberg when you test many hypotheses and can accept a controlled share of false discoveries in exchange for more power.

Holm controls the same family-wise error rate as Bonferroni but is uniformly more powerful, so it rejects at least as many hypotheses. There is no reason to prefer plain Bonferroni.

Sort the m p-values, multiply the i-th smallest by m/i, then take a running minimum from the largest down so the adjusted values never decrease, capping them at 1.

Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page.