When to use it
Use it whenever you report an average of per-item scores: LLM-as-judge ratings (1–10, 1–5), human preference scores, per-example metric scores (sentence-level BLEU, ROUGE, cosine similarity), or latencies. With one column it gives the mean with a confidence interval. With two columns it tests whether system A scores higher than system B on the same items.
Why pairing matters
Some prompts are easy and score high for every system; others are hard for all. That shared variation inflates each system’s own interval but cancels out in the per-item difference. In the example, the two systems’ 95% intervals (5.84–7.06 and 5.04–6.31) overlap, yet the paired difference of 0.78 points has an interval of 0.33 to 1.22 and p = 0.001. Comparing two separate intervals would have wrongly concluded “no difference”.
t-interval or bootstrap?
The t-interval (and paired t-test) assumes the mean is approximately normal, which holds well for a few dozen items or more even when individual scores are lumpy. The percentile bootstrap resamples items and makes no distributional assumption. We show both; when they agree, as in the example, you can report either. Our t-intervals and paired t-test match SciPy to machine precision.
Common mistakes
- Comparing overlapping intervals instead of testing the paired difference.
- Treating repeated judge samples as separate items. Average them per item first.
- Averaging corpus-level metrics. Corpus BLEU is not the mean of sentence BLEU; resample whole items and recompute instead.
- Ignoring judge reliability. An unreliable judge adds noise; check it with Cohen’s kappa.
Related
Frequently asked questions
Use the t-interval: mean plus or minus a t critical value times the standard error (SD divided by the square root of the number of items). A bootstrap over items gives a distribution-free check.
If both were scored on the same items, use a paired t-test on the per-item differences. It is far more sensitive than comparing two separate confidence intervals.
Each interval includes item-to-item variation that both systems share. The paired difference removes it, so it can be significant even when the separate intervals overlap.
Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.