Interpreting Effect Size for Better Data Storytelling
A p-value tells your reader whether an effect is distinguishable from noise; an effect size tells them how big it is — and that is the part of the story they actually care about. The trick is translating effect sizes into terms people can picture. A Cohen’s d of 0.5 means little to most readers; “the two groups overlap by 80%, but a typical treated person beats a typical untreated one 64% of the time” means a lot.
See it first
Drag Cohen’s d below and watch the two distributions separate. Try the quiz to test your intuition.
Open the effect size visualizer full screen.
Four ways to say the same thing
For two normal groups with equal spread, every Cohen’s d corresponds to these plain-language quantities:
| Cohen’s d | Treatment above control mean (U3) | Overlap | Random treated beats random control | Number needed to treat |
|---|---|---|---|---|
| 0.2 | 57.9% | 92.0% | 55.6% | 8.9 |
| 0.5 | 69.1% | 80.3% | 63.8% | 3.6 |
| 0.8 | 78.8% | 68.9% | 71.4% | 2.3 |
Pick the one that fits your audience. Teachers relate to “percent of students above the old average”; clinicians to the number needed to treat; general readers to “how often does one beat the other”.
Why “small, medium, large” mislead
Cohen offered 0.2, 0.5 and 0.8 as fallback benchmarks for when nothing better was known. Context decides importance: a d of 0.2 on survival or test scores across a whole school system can be enormous, while a d of 0.8 on a lab reaction-time task may change nothing that matters. Compare your effect with others in your field, with the cost of the intervention, and with what is practically meaningful.
Always pair it with uncertainty
An effect size without a confidence interval tells half the story. “d = 0.45” from 20 people and from 2,000 people are very different claims. Report the interval, and say what its ends would mean: a range from “negligible” to “large” is honest and informative. Our effect size calculator and mean score calculator give intervals directly.
A storytelling template
Suppose a new teaching method raises test scores by d = 0.45. A sentence that works for both experts and lay readers:
Students taught with the new method scored 0.45 standard deviations higher [add the confidence interval]. In practical terms, about 67% of them scored above the average student taught the old way, and a randomly chosen student from the new group outscored one from the old group about 62% of the time.
Those translations come straight from d: Φ(0.45) = 67.4% and Φ(0.45/√2) = 62.5%.
Effect sizes for other kinds of data
- Yes/no outcomes: report the risk difference, relative risk or odds ratio with its interval, and the number needed to treat; the odds ratio calculator gives all of them.
- Relationships: a correlation r, or R² for the share of variance explained — remembering that r = 0.3 explains only 9%.
- Several groups: η² from an ANOVA, plus the pairwise differences that matter.
- Categorical tables: Cramér’s V alongside the chi-square test.
Common mistakes
- Calling a significant result “large”. With big samples, tiny effects are significant.
- Reporting only the label. “A medium effect” without the number and interval hides the evidence.
- Mixing up d and r. They live on different scales; d = 0.5 is roughly r = 0.24.