An eval score is an estimate, so it needs an error bar. If a model answers 412 of 500 items correctly (82.4%), the 95% Wilson confidence interval is 78.8% to 85.5% — roughly ±3.3 points. Use the Wilson interval, not the textbook “±1.96×SE” (Wald) interval, which breaks near 0% or 100%: at 100/100 correct, Wald claims zero uncertainty while Wilson gives 96.3%–100%. To tell whether two models really differ, use a paired McNemar test rather than comparing their intervals.

Why an eval score needs a confidence interval

A benchmark is a sample of all the questions a model could face. Score the same model on a different 500 items and you would get a slightly different number. The confidence interval says how far the score could plausibly move from sampling alone. Without it, a 1–2 point “improvement” on a few hundred items is often noise, not progress.

Why Wilson instead of ±1.96 × SE

The familiar Wald interval, p ± z√(p(1−p)/n), undercovers badly when accuracy is near 0% or 100% — exactly where strong models live. At 98/100 correct, Wald gives 95.3%–100%, while Wilson gives a more honest 93.0%–99.4%. At 100/100, Wald collapses to [100%, 100%], claiming certainty. The Wilson score interval keeps close to its nominal coverage at every accuracy level and is the default in statsmodels and most modern eval tooling.

How many items you need

The interval narrows with the square root of the number of items, so halving the margin takes four times the data. At about 82% accuracy, a ±5-point margin needs about 223 items, ±3 points about 620, ±2 points about 1,393, and ±1 point about 5,572. If you have no idea of the accuracy yet, plan for the worst case (50%): 385, 1,068, 2,401 and 9,604 items. The calculator prints this table for your own numbers.

Common mistakes

  • Comparing two models by whether their intervals overlap. Overlapping intervals can still hide a real difference, because both models answered the same items. Use the McNemar test.
  • Treating repeated samples of the same item as independent. If you sample each prompt 5 times, n is the number of prompts, not prompts × 5 — otherwise the interval is far too narrow.
  • Ignoring that the test set is not the deployment distribution. The interval covers sampling noise only; it says nothing about whether the benchmark matches real usage.

Related

AI eval hubEval confidence intervalCompare two modelsEval sample sizepass@kCohen's kappa

Frequently asked questions

Use the Wilson score interval on the number of correct answers out of the total items. For 412 correct out of 500 the 95% interval is about 78.8% to 85.5%.

The Wald interval (accuracy plus or minus 1.96 standard errors) undercovers near 0% and 100% and collapses to zero width at perfect scores. The Wilson interval stays accurate across the whole range.

About 1,400 items at 80-85% accuracy and up to 2,401 items in the worst case of 50% accuracy, at 95% confidence.

Computed in your browser; formulas verified against SciPy and statsmodels. See the methodology page and the AI & LLM evaluation hub.