Krippendorff’s alpha measures how reliably several raters agree, beyond chance, and it tolerates missing ratings and any measurement scale. 1 is perfect reliability and 0 is chance level; Krippendorff suggests α ≥ 0.800 for firm conclusions and ≥ 0.667 for tentative ones. The measurement level matters: in our example, four annotators score 10 items on a 1–5 scale with some blanks. Treated as ordinal, α = 0.90 (95% bootstrap CI 0.70–0.93). Treated as nominal, where a 4 vs 5 counts as a full disagreement, the same data gives α = 0.52.

Separate with commas or tabs (paste from a spreadsheet). Leave a cell empty, or write NA, for a missing rating.

Why use Krippendorff’s alpha

It is the most general agreement coefficient. It works with any number of raters, tolerates missing ratings (so not every annotator has to label every item), and supports nominal, ordinal, interval and ratio data with the right notion of “how far apart” two ratings are. That makes it the usual choice for annotation projects in content analysis and NLP, and for checking human labels before using them to evaluate a model.

Pick the right level of measurement

  • Nominal — unordered labels (spam / not spam, topic). Any difference is a full disagreement.
  • Ordinal — ordered ranks (1–5 quality scores, Likert items). Near misses count less, based on how many ratings fall between the two values.
  • Interval — numbers where differences are meaningful; disagreement grows with the squared difference.
  • Ratio — numbers with a true zero (counts, durations); differences are judged relative to size.

In the example, the four annotators rarely differ by more than one point, so ordinal α (0.90) is far higher than nominal α (0.52). Choosing nominal for a rating scale understates reliability; choosing interval for unordered labels is meaningless.

How it works

α = 1 − Do / De: one minus the ratio of observed disagreement to the disagreement expected by chance, both computed from a coincidence matrix of all pairable ratings within items. Our implementation matches the reference krippendorff Python package to machine precision at all four levels, including with missing data.

Common mistakes

  • Dropping every item with a missing rating. Alpha uses all items with at least two ratings; throwing data away only widens the interval.
  • Reporting alpha without the level of measurement — the same data can give very different values.
  • Ignoring uncertainty. With 10 items, as here, the 95% interval for ordinal α still runs from 0.70 to 0.93.

Related

AI eval hubCohen's kappaFleiss' kappaKrippendorff's alphaArena Elo ratingsF1 confidence intervalCompare two models

Frequently asked questions

Krippendorff recommends alpha of at least 0.800 for reliable data, and 0.667 as the lowest value for tentative conclusions.

Yes. It uses every item that has at least two ratings, so raters do not need to label every item.

Nominal for unordered categories, ordinal for ordered scores like 1 to 5, interval for numeric data where differences matter, and ratio for numeric data with a true zero.

Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.