Arena-style leaderboards turn head-to-head votes (“which answer is better?”) into ratings with the Bradley–Terry model, shown on the familiar Elo scale. A 100-point gap means the higher model wins about 64% of the time; 400 points means about 91%. Enter win counts for each pair of models. Example: four models, 600 comparisons: Model A rates 1119 ± 33, B 1020 ± 31, C 972 ± 31, D 888 ± 33, and A is expected to beat B 63.9% of the time. The calculator also tests each gap between neighbouring ranks, because a leaderboard order is only meaningful where the confidence intervals separate.

Commas or tabs. Ties count as half a win for each side. Repeated pairs are added together.

From pairwise votes to a leaderboard

In an arena evaluation, people (or a judge model) see two answers to the same prompt and pick the better one. The Bradley–Terry model assumes each model has a strength s, and model i beats model j with probability si / (si + sj). We find the strengths that best explain all the observed wins at once, then put them on the Elo scale: rating = 1000 + 400·log10(s), centred so the average model is 1000.

Why not update Elo game by game?

Classic online Elo updates ratings after each game, so the result depends on the order games were played and on an arbitrary K-factor. For a fixed set of model comparisons, that order carries no meaning. The Bradley–Terry fit uses all games simultaneously, gives the same answer in any order, and comes with standard errors. It is the model behind arena-style LLM leaderboards.

Reading the confidence intervals

Each rating’s interval reflects how many comparisons involve that model and how decisive they were. The rank order is only established where neighbouring models are statistically separated, which the “gap to next” column tests directly using the covariance of the two ratings. Two models whose intervals overlap heavily should be reported as tied. In the example, A leads B by 99 points (p = 0.0002) and C leads D by 84 (p = 0.0013), but B’s 48-point lead over C is not significant (p = 0.06) — 600 comparisons are not enough to order those two.

Requirements and common mistakes

  • The comparisons must connect every model. A model that never loses (or never wins) has no finite rating; the calculator tells you when the data cannot identify the ratings.
  • Ratings are relative. 1000 is the average of the models you entered, so ratings from different pools are not comparable.
  • Unbalanced prompts bias results. If one model faces only hard prompts, its rating suffers; randomise prompt and opponent assignment.
  • Position bias in judges. LLM judges often favour the first answer shown; randomise the order and use Cohen’s kappa to check judges against humans.

Related

AI eval hubCohen's kappaFleiss' kappaKrippendorff's alphaArena Elo ratingsF1 confidence intervalCompare two modelsCompare two models (McNemar)

Frequently asked questions

They fit a Bradley-Terry model to the pairwise votes and show the strengths on the Elo scale, where a 400-point gap corresponds to 10-to-1 odds of winning.

The higher-rated model is expected to win about 64% of head-to-head comparisons. A 200-point gap means about 76%, and 400 points about 91%.

Only if its rating is statistically separated from the next model. Check whether the confidence intervals overlap, or use the gap test in this calculator.

Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.