Why kappa instead of percent agreement
If 90% of answers are “pass”, two raters who both say “pass” almost every time will agree around 80% of the time by luck alone. Kappa subtracts that expected agreement: κ = (po − pe) / (1 − pe), where po is observed agreement and pe is the agreement expected from each rater’s label frequencies.
Validating an LLM judge
Before trusting an LLM-as-judge to grade outputs at scale, have humans label a sample (100–300 items is typical) and compute kappa between the human labels and the judge. Compare it to human–human kappa on the same items: a judge that agrees with humans about as well as humans agree with each other is usually good enough. Report the confidence interval — on 20 items, as in the example, it is very wide (0.35–1.00).
Weighted kappa for 1–5 scores
For ordered ratings, a 4 vs 5 disagreement is milder than a 1 vs 5. Weighted kappa gives partial credit: linear weights penalise disagreements in proportion to distance, quadratic weights by squared distance (quadratic weighted kappa is common in ML competitions). On a 15-item 1–5 example, unweighted κ is 0.58 but quadratic weighted κ is 0.90, because most disagreements are off by one.
Interpreting the value
Landis & Koch’s widely used labels: below 0 poor, 0–0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, above 0.80 almost perfect. Treat them as rough guides, not thresholds.
Common mistakes
- Reporting only percent agreement, which rewards raters for agreeing on the majority class.
- Using Cohen’s kappa for more than two raters. Use Fleiss’ kappa or Krippendorff’s alpha instead.
- Misaligned items. Both label lists must be in the same item order.
Related
Frequently asked questions
By the Landis and Koch guide, 0.61 to 0.80 is substantial and above 0.80 almost perfect agreement. For an LLM judge, compare its kappa with humans to the kappa between two humans on the same items.
Have humans label a sample of items, run the LLM judge on the same items, and compute Cohen's kappa between the two label lists. Use weighted kappa for ordered scores like 1 to 5.
Cohen's kappa treats every disagreement equally. Weighted kappa gives partial credit to near misses on ordered scales, using linear or quadratic weights.
Computed in your browser; formulas verified against SciPy and statsmodels. See the methodology page and the AI & LLM evaluation hub.