How big does n need to be?
The popular “n ≥ 30” rule is only a rough guide. What matters is how skewed the population is: the skewness of the sample mean shrinks like γ/√n, where γ is the population’s skewness. For the Skewed (exponential) population here, γ = 2:
| Sample size n | 1 | 5 | 10 | 30 | 100 |
|---|---|---|---|---|---|
| Skewness of the mean (2/√n) | 2.00 | 0.89 | 0.63 | 0.37 | 0.20 |
| Spread of the mean (5/√n) | 5.00 | 2.24 | 1.58 | 0.91 | 0.50 |
Symmetric populations — Uniform, Bimodal, Die roll — have γ = 0, so their means look normal by n = 5–10. In 20,000 simulated samples per setting, the share of means within ±1 SE of the true mean was 87% for skewed data at n = 1 (far from the normal 68%) but 68% by n = 30. The readouts above let you check this yourself.
When the Central Limit Theorem fails
The theorem needs a finite variance. The Heavy-tailed option draws from a Cauchy distribution, which has no finite mean or variance: occasional values are so extreme that they dominate any average. The mean of n Cauchy values is itself Cauchy with the same spread, so no bell ever forms. In our simulation the middle 50% of sample means stayed about 2 units wide at n = 1, 10 and 100, and about 6% of means landed outside the plotted window every time. Real data with very heavy tails — some financial returns, file sizes, insurance claims — can behave like this, which is why medians and robust methods exist.
Why it matters
The CLT is why confidence intervals and t-tests work for data that is not normal: they rely on the distribution of the mean. See it in action in the confidence interval simulator, or read the step-by-step CLT walkthrough on the blog.