When to use Fleiss’ kappa
Use it when every item is rated by the same number of raters (three or more) into categories — for example three annotators labelling sentiment, or several LLM judges grading answers as good/okay/bad. The raters do not have to be the same people for every item. For exactly two raters use Cohen’s kappa; with missing ratings or ordered scales, Krippendorff’s alpha is the more flexible choice.
How it is calculated
For each item, Pi is the share of rater pairs that agree. Their average, P̄, is the observed agreement. P̄e is the agreement expected if raters picked categories at their overall rates: the sum of squared category shares. Then
κ = (P̄ − P̄e) / (1 − P̄e)
Our implementation matches statsmodels to machine precision and reproduces the textbook example on Wikipedia (κ = 0.210).
Using it to check a panel of LLM judges
If you grade outputs with several judge models or prompts, Fleiss’ kappa tells you whether they are measuring the same thing. Low agreement means the rubric is ambiguous or the judges are unreliable, so averaging their scores will mostly average noise. Fix the rubric before scaling up, and compare against humans with Cohen’s kappa.
Common mistakes
- Reporting kappa without an interval. On a few dozen items the uncertainty is huge; the example’s 95% interval runs from 0.14 to 0.74.
- Treating ordered scores as unordered. For 1–5 ratings, Fleiss’ kappa ignores that 4 vs 5 is a near miss; Krippendorff’s alpha (ordinal) does not.
- Comparing kappas across datasets with different category balance. Kappa depends on how common each category is, so the same rater quality can produce different values.
Related
Frequently asked questions
It measures agreement among three or more raters who each assign items to categories, correcting for agreement expected by chance. It is used for annotation quality and for checking agreement among several LLM judges.
Cohen's kappa handles exactly two raters. Fleiss' kappa handles any fixed number of raters per item, and the raters may differ from item to item.
Fleiss' kappa needs the same number of ratings for every item. With missing ratings use Krippendorff's alpha, which handles them directly.
Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.