An average score, such as a mean LLM-judge rating out of 10, needs a confidence interval, and comparing two systems on the same items needs a paired test. Paste one score per item, or two (system A, system B) for a comparison. Example: on 40 prompts, system A averages 6.45 (95% CI 5.84–7.06) and system B 5.68 (5.04–6.31). The two intervals overlap, yet the paired difference, +0.78 (95% CI 0.33–1.22, p = 0.001), is clearly significant: pairing removes the prompt-to-prompt variation that both systems share.

Commas or tabs. Any numeric scale works (1–10 judge ratings, 0–1 metric scores, latencies). For two columns, both systems must be scored on the same items in the same order.

When to use it

Use it whenever you report an average of per-item scores: LLM-as-judge ratings (1–10, 1–5), human preference scores, per-example metric scores (sentence-level BLEU, ROUGE, cosine similarity), or latencies. With one column it gives the mean with a confidence interval. With two columns it tests whether system A scores higher than system B on the same items.

Why pairing matters

Some prompts are easy and score high for every system; others are hard for all. That shared variation inflates each system’s own interval but cancels out in the per-item difference. In the example, the two systems’ 95% intervals (5.84–7.06 and 5.04–6.31) overlap, yet the paired difference of 0.78 points has an interval of 0.33 to 1.22 and p = 0.001. Comparing two separate intervals would have wrongly concluded “no difference”.

t-interval or bootstrap?

The t-interval (and paired t-test) assumes the mean is approximately normal, which holds well for a few dozen items or more even when individual scores are lumpy. The percentile bootstrap resamples items and makes no distributional assumption. We show both; when they agree, as in the example, you can report either. Our t-intervals and paired t-test match SciPy to machine precision.

Common mistakes

  • Comparing overlapping intervals instead of testing the paired difference.
  • Treating repeated judge samples as separate items. Average them per item first.
  • Averaging corpus-level metrics. Corpus BLEU is not the mean of sentence BLEU; resample whole items and recompute instead.
  • Ignoring judge reliability. An unreliable judge adds noise; check it with Cohen’s kappa.

Related

AI eval hubCohen's kappaFleiss' kappaKrippendorff's alphaArena Elo ratingsF1 confidence intervalCompare two modelsPaired bootstrap testT-test

Frequently asked questions

Use the t-interval: mean plus or minus a t critical value times the standard error (SD divided by the square root of the number of items). A bootstrap over items gives a distribution-free check.

If both were scored on the same items, use a paired t-test on the per-item differences. It is far more sensitive than comparing two separate confidence intervals.

Each interval includes item-to-item variation that both systems share. The paired difference removes it, so it can be significant even when the separate intervals overlap.

Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.