Fleiss’ kappa measures agreement among three or more raters beyond what chance would produce. It extends Cohen’s kappa (which handles exactly two raters) to any fixed number of raters per item. 0 means chance-level agreement and 1 means perfect agreement. Example: three LLM judges grade 12 answers as good/okay/bad and agree on 66.7% of rater pairs, against 35.3% expected by chance, so κ = 0.48 (“moderate”, p < 0.0001). The 95% bootstrap interval, 0.14–0.74, shows how little 12 items can pin down. With missing ratings or ordered scores, use Krippendorff’s alpha instead.

Separate labels with commas or tabs (you can paste straight from a spreadsheet). Every line needs the same number of ratings.

When to use Fleiss’ kappa

Use it when every item is rated by the same number of raters (three or more) into categories — for example three annotators labelling sentiment, or several LLM judges grading answers as good/okay/bad. The raters do not have to be the same people for every item. For exactly two raters use Cohen’s kappa; with missing ratings or ordered scales, Krippendorff’s alpha is the more flexible choice.

How it is calculated

For each item, Pi is the share of rater pairs that agree. Their average, P̄, is the observed agreement. P̄e is the agreement expected if raters picked categories at their overall rates: the sum of squared category shares. Then

κ = (P̄ − P̄e) / (1 − P̄e)

Our implementation matches statsmodels to machine precision and reproduces the textbook example on Wikipedia (κ = 0.210).

Using it to check a panel of LLM judges

If you grade outputs with several judge models or prompts, Fleiss’ kappa tells you whether they are measuring the same thing. Low agreement means the rubric is ambiguous or the judges are unreliable, so averaging their scores will mostly average noise. Fix the rubric before scaling up, and compare against humans with Cohen’s kappa.

Common mistakes

  • Reporting kappa without an interval. On a few dozen items the uncertainty is huge; the example’s 95% interval runs from 0.14 to 0.74.
  • Treating ordered scores as unordered. For 1–5 ratings, Fleiss’ kappa ignores that 4 vs 5 is a near miss; Krippendorff’s alpha (ordinal) does not.
  • Comparing kappas across datasets with different category balance. Kappa depends on how common each category is, so the same rater quality can produce different values.

Related

AI eval hubCohen's kappaFleiss' kappaKrippendorff's alphaArena Elo ratingsF1 confidence intervalCompare two models

Frequently asked questions

It measures agreement among three or more raters who each assign items to categories, correcting for agreement expected by chance. It is used for annotation quality and for checking agreement among several LLM judges.

Cohen's kappa handles exactly two raters. Fleiss' kappa handles any fixed number of raters per item, and the raters may differ from item to item.

Fleiss' kappa needs the same number of ratings for every item. With missing ratings use Krippendorff's alpha, which handles them directly.

Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.