Calculators and a short guide for reporting model evals with real error bars.
Written & reviewed by Dackohn; last reviewed October 2026.
Most model comparisons skip the statistics, so many reported gains are noise. Five tools fix that: put a Wilson confidence interval on every accuracy score; compare two models with a paired McNemar test on the same items; size the eval before you run it; estimate pass@k with the unbiased estimator; and check an LLM judge against humans with Cohen’s kappa. Example: 412/500 correct is 82.4% with a 95% interval of 78.8%–85.5% — a 2-point lead over another model on the same 500 items is often not significant.
| Question | Calculator | What it gives you |
|---|---|---|
| How precise is my accuracy score? | Eval confidence interval | Wilson interval + items needed for a target margin |
| Is model A really better than model B? | Compare two models (McNemar) | Exact paired test + CI for the gap |
| How many eval items do I need? | Eval sample size | Paired sample size from accuracies + disagreement |
| What is pass@k for my code eval? | pass@k | Unbiased estimator from the Codex paper |
| Can I trust my LLM judge? | Cohen's kappa | Kappa, CI, and weighted kappa for 1–5 scores |
Two models scored on the same items share the items’ difficulty. Easy items are right for both and broken items wrong for both, and neither tells you which model is better. A paired test keeps only the disagreements, which is why it detects real differences with far fewer items than comparing two independent scores. In our sample-size calculator, an 80% vs 75% comparison needs 1,091 items per model unpaired, but only 469 shared items when the models disagree on 15% of them.
Report accuracy with a Wilson confidence interval, compare models with a paired McNemar test on the same items, size the eval in advance for the gap you care about, use the unbiased pass@k estimator for sampled outputs, and validate any LLM judge against humans with Cohen's kappa.
For a single score, about 1,068 items gives a margin of plus or minus 3 points in the worst case. For comparing two models, the number depends on the gap and how often the models disagree, often a few hundred to a few thousand items.
These tools build on general methods covered elsewhere on the site: confidence intervals, two-proportion tests, sample size and Type I & II errors. All calculations run in your browser and are verified against SciPy and statsmodels (methodology).