Why use Krippendorff’s alpha
It is the most general agreement coefficient. It works with any number of raters, tolerates missing ratings (so not every annotator has to label every item), and supports nominal, ordinal, interval and ratio data with the right notion of “how far apart” two ratings are. That makes it the usual choice for annotation projects in content analysis and NLP, and for checking human labels before using them to evaluate a model.
Pick the right level of measurement
- Nominal — unordered labels (spam / not spam, topic). Any difference is a full disagreement.
- Ordinal — ordered ranks (1–5 quality scores, Likert items). Near misses count less, based on how many ratings fall between the two values.
- Interval — numbers where differences are meaningful; disagreement grows with the squared difference.
- Ratio — numbers with a true zero (counts, durations); differences are judged relative to size.
In the example, the four annotators rarely differ by more than one point, so ordinal α (0.90) is far higher than nominal α (0.52). Choosing nominal for a rating scale understates reliability; choosing interval for unordered labels is meaningless.
How it works
α = 1 − Do / De: one minus the ratio of observed disagreement to the disagreement expected by chance, both computed from a coincidence matrix of all pairable ratings within items. Our implementation matches the reference krippendorff Python package to machine precision at all four levels, including with missing data.
Common mistakes
- Dropping every item with a missing rating. Alpha uses all items with at least two ratings; throwing data away only widens the interval.
- Reporting alpha without the level of measurement — the same data can give very different values.
- Ignoring uncertainty. With 10 items, as here, the 95% interval for ordinal α still runs from 0.70 to 0.93.
Related
Frequently asked questions
Krippendorff recommends alpha of at least 0.800 for reliable data, and 0.667 as the lowest value for tentative conclusions.
Yes. It uses every item that has at least two ratings, so raters do not need to label every item.
Nominal for unordered categories, ordinal for ordered scores like 1 to 5, interval for numeric data where differences matter, and ratio for numeric data with a true zero.
Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.