Written & reviewed by Dackohn; last reviewed October 2026.

Most model comparisons skip the statistics, so many reported gains are noise. Five tools fix that: put a Wilson confidence interval on every accuracy score; compare two models with a paired McNemar test on the same items; size the eval before you run it; estimate pass@k with the unbiased estimator; and check an LLM judge against humans with Cohen’s kappa. Example: 412/500 correct is 82.4% with a 95% interval of 78.8%–85.5% — a 2-point lead over another model on the same 500 items is often not significant.

Which calculator answers which question

QuestionCalculatorWhat it gives you
How precise is my accuracy score?Eval confidence intervalWilson interval + items needed for a target margin
Is model A really better than model B?Compare two models (McNemar)Exact paired test + CI for the gap
How many eval items do I need?Eval sample sizePaired sample size from accuracies + disagreement
What is pass@k for my code eval?pass@kUnbiased estimator from the Codex paper
Can I trust my LLM judge?Cohen's kappaKappa, CI, and weighted kappa for 1–5 scores

A reporting checklist for model evals

  1. State n. Report the number of items, not just the percentage.
  2. Add an interval. Give each score a 95% Wilson interval: 82.4% [78.8, 85.5].
  3. Test comparisons with pairing. If two models saw the same items, use McNemar’s test and report the p-value and the interval for the gap.
  4. Count independent units correctly. Repeated samples of one prompt are one item, not several; average them per item first.
  5. Correct for many comparisons. When testing many prompts or checkpoints, adjust α or confirm the winner on a fresh set.
  6. Plan the size. Decide the smallest gap that matters and size the eval for it before running.
  7. Validate automated graders. Report kappa between your LLM judge and human labels, alongside human–human kappa.

Why paired comparisons matter so much

Two models scored on the same items share the items’ difficulty. Easy items are right for both and broken items wrong for both, and neither tells you which model is better. A paired test keeps only the disagreements, which is why it detects real differences with far fewer items than comparing two independent scores. In our sample-size calculator, an 80% vs 75% comparison needs 1,091 items per model unpaired, but only 469 shared items when the models disagree on 15% of them.

Frequently asked questions

Report accuracy with a Wilson confidence interval, compare models with a paired McNemar test on the same items, size the eval in advance for the gap you care about, use the unbiased pass@k estimator for sampled outputs, and validate any LLM judge against humans with Cohen's kappa.

For a single score, about 1,068 items gives a margin of plus or minus 3 points in the worst case. For comparing two models, the number depends on the gap and how often the models disagree, often a few hundred to a few thousand items.

These tools build on general methods covered elsewhere on the site: confidence intervals, two-proportion tests, sample size and Type I & II errors. All calculations run in your browser and are verified against SciPy and statsmodels (methodology).