What calibration means
A classifier or LLM that says “90% sure” should be right about 90% of the time. Calibration matters whenever confidence scores drive decisions: abstaining or escalating to a human below a threshold, ranking answers, or combining models. Modern neural networks, including LLMs after fine-tuning, are often over-confident.
How ECE is computed
Group predictions into equal-width confidence bins. In each bin, compare the mean confidence with the actual accuracy. ECE is the average absolute gap, weighted by the share of predictions in each bin; MCE is the largest gap in any bin. The Brier score is the mean squared difference between confidence and the 0/1 outcome, and rewards both calibration and sharpness. Our per-bin values match scikit-learn’s calibration_curve and our Brier score matches brier_score_loss.
Reading the reliability diagram
Each bar is the accuracy within one confidence bin, the orange dot is the model’s mean confidence in that bin, and the dashed diagonal is perfect calibration. Bars below their dots mean over-confidence; bars above mean under-confidence. Hover or tab through the bins for exact values.
Common mistakes
- Comparing ECE across different bin counts or test sizes. ECE is biased upward on small samples and depends on the number of bins, so report both and include the interval.
- Judging calibration by accuracy alone. A 90%-accurate model can still be badly calibrated.
- Using too few predictions. With a few dozen predictions per bin the gaps are mostly noise; widen the bins or collect more data.
Related
Frequently asked questions
ECE measures how far a model's confidence is from its actual accuracy. Predictions are grouped into confidence bins, and the absolute gap between mean confidence and accuracy in each bin is averaged, weighted by bin size.
Lower is better; 0 means perfect calibration. Values below about 0.02 to 0.05 are usually considered well calibrated, but ECE depends on the number of bins and the sample size, so report them alongside it.
ECE measures only the confidence-accuracy gap. The Brier score is the mean squared error of the probabilities and combines calibration with how decisive the predictions are.
Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page and the AI & LLM Evaluation hub.