What pass@k measures
For code generation and many reasoning tasks there is no single right output to compare against; instead you check whether a generated answer passes the tests. pass@k asks: if the model gets k attempts at a problem, how likely is it that at least one passes? pass@1 is the expected single-attempt success rate; pass@10 or pass@100 describe a setting where a user or system can try several samples and keep one that works.
Why the unbiased estimator
The obvious approach — sample exactly k answers and check whether any passes — has high variance. The Codex paper instead draws n ≥ k samples, counts the c correct ones, and computes the probability that a random subset of k contains at least one correct sample: 1 − C(n−c, k)/C(n, k). This is unbiased. Plugging the sample success rate into 1 − (1 − c/n)k looks reasonable but is biased, especially for large k. We compute the estimator with the numerically stable product form, so large n is fine.
Worked example
One problem, n = 200 samples, c = 15 correct: pass@1 = 7.50%, pass@5 = 32.56%, pass@10 = 55.00%, pass@50 = 98.89%. For a benchmark, enter one line per problem; the tool averages pass@k across them, which is how HumanEval-style results are reported.
Common mistakes
- Using n = k. It is unbiased but very noisy; use n well above k (the Codex paper used n = 200).
- Comparing pass@k across different temperatures without saying so. Higher temperature lowers pass@1 but often raises pass@100; report the sampling settings.
- Forgetting uncertainty. pass@k is still an average over a finite set of problems; with a few dozen problems, differences of a few points are often noise.
Related
Frequently asked questions
Draw n samples per problem (n at least k), count c correct ones, and compute 1 - C(n-c, k) / C(n, k). Average that across problems. This is the unbiased estimator from the Codex paper (Chen et al., 2021).
Plugging the sample success rate into that formula gives a biased estimate of pass@k. The combinatorial estimator is unbiased.
At least k, and preferably many more. The Codex paper used 200 samples per problem to estimate pass@1, pass@10 and pass@100 with low variance.
Computed in your browser; formulas verified against SciPy and statsmodels. See the methodology page and the AI & LLM evaluation hub.