Cohen’s kappa measures how much two raters agree beyond what chance alone would produce. It runs from −1 to 1: 0 means chance-level agreement and 1 means perfect agreement. Raw percent agreement is misleading when one label dominates. Example: a human and an LLM judge rate 20 answers pass/fail and agree on 85%, but 53% agreement was expected by chance, so κ = 0.68 (“substantial”, 95% CI 0.35–1.00, p = 0.002). For ordered scores such as 1–5 ratings, use weighted kappa, which gives partial credit for near misses.

Separate labels with commas or new lines, in the same item order for both raters. Numeric labels (e.g. 1–5) also get weighted kappa.

Why kappa instead of percent agreement

If 90% of answers are “pass”, two raters who both say “pass” almost every time will agree around 80% of the time by luck alone. Kappa subtracts that expected agreement: κ = (po − pe) / (1 − pe), where po is observed agreement and pe is the agreement expected from each rater’s label frequencies.

Validating an LLM judge

Before trusting an LLM-as-judge to grade outputs at scale, have humans label a sample (100–300 items is typical) and compute kappa between the human labels and the judge. Compare it to human–human kappa on the same items: a judge that agrees with humans about as well as humans agree with each other is usually good enough. Report the confidence interval — on 20 items, as in the example, it is very wide (0.35–1.00).

Weighted kappa for 1–5 scores

For ordered ratings, a 4 vs 5 disagreement is milder than a 1 vs 5. Weighted kappa gives partial credit: linear weights penalise disagreements in proportion to distance, quadratic weights by squared distance (quadratic weighted kappa is common in ML competitions). On a 15-item 1–5 example, unweighted κ is 0.58 but quadratic weighted κ is 0.90, because most disagreements are off by one.

Interpreting the value

Landis & Koch’s widely used labels: below 0 poor, 0–0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, above 0.80 almost perfect. Treat them as rough guides, not thresholds.

Common mistakes

  • Reporting only percent agreement, which rewards raters for agreeing on the majority class.
  • Using Cohen’s kappa for more than two raters. Use Fleiss’ kappa or Krippendorff’s alpha instead.
  • Misaligned items. Both label lists must be in the same item order.

Related

AI eval hubEval confidence intervalCompare two modelsEval sample sizepass@kCohen's kappaChi-square test

Frequently asked questions

By the Landis and Koch guide, 0.61 to 0.80 is substantial and above 0.80 almost perfect agreement. For an LLM judge, compare its kappa with humans to the kappa between two humans on the same items.

Have humans label a sample of items, run the LLM judge on the same items, and compute Cohen's kappa between the two label lists. Use weighted kappa for ordered scores like 1 to 5.

Cohen's kappa treats every disagreement equally. Weighted kappa gives partial credit to near misses on ordered scales, using linear or quadratic weights.

Computed in your browser; formulas verified against SciPy and statsmodels. See the methodology page and the AI & LLM evaluation hub.