What Makes a Free Online Statistics Tool Reliable and Accurate?
A statistics tool is reliable when it uses the right method for the situation, matches established reference software on known examples, behaves sensibly at the edges, and tells you exactly what it computed. Arithmetic is rarely the problem; choosing the method is. This post explains how we check the calculators on this site, with real numbers — including the mistakes the checks caught — and ends with a checklist you can apply to any free tool.
1. The method matters more than the arithmetic
Two calculators can both compute flawlessly and still disagree, because they use different methods. Two examples:
Confidence intervals for a proportion. The textbook “Wald” interval, p ± 1.96√(p(1−p)/n), is still the default in many tools. We computed its exact coverage — how often a 95% interval actually contains the true value — by summing binomial probabilities:
| Sample size, true proportion | Wald coverage | Wilson coverage |
|---|---|---|
| n = 20, p = 0.50 | 95.9% | 95.9% |
| n = 20, p = 0.90 | 87.6% | 95.7% |
| n = 20, p = 0.95 | 63.9% | 92.5% |
| n = 100, p = 0.98 | 86.6% | 94.9% |
Near 0% or 100% the “95%” Wald interval can be wrong a third of the time; at 100 successes out of 100 it collapses to zero width. That is why our proportion intervals use Wilson.
Comparing two means. Student’s classic t-test assumes both groups have the same variance; Welch’s version does not. In one test case — a group of 8 with a small spread against a group of 25 with a large one — Welch gave p = 0.023 and the pooled test p = 0.130: opposite conclusions from the same data. Welch is the safer default, costing almost nothing when the variances happen to be equal.
2. Check every result against reference software
We test each calculation against established scientific libraries — SciPy, statsmodels, scikit-learn and specialist packages — on hundreds of random inputs, not just one example. The browser-based tools are tested by running their actual JavaScript in a separate engine and comparing outputs. Typical worst-case differences:
- Wilson intervals vs statsmodels: within 5×10−10
- Cohen’s kappa, its standard error and interval vs statsmodels: within 5×10−8
- Krippendorff’s alpha (four measurement levels, missing data) vs the reference package: within 4×10−16
- Arena-style Bradley–Terry ratings and their standard errors vs a statsmodels logistic regression: within 4×10−6
The point is not the decimals but the breadth: a tool that matches on random inputs, edge cases included, is far more trustworthy than one that matches a single textbook example.
3. Test the claims, not just the formulas
Some promises can only be checked by simulation. Our sample-size calculator promises 80% power; in 20,000 simulated studies at the recommended size, the test detected the effect 79.7–80.1% of the time. The sequential A/B test promises that checking results every day will not inflate false positives; in 4,000 simulated tests with no real difference, peeking 100 times produced false positives 0.6–2.1% of the time — while a standard test checked the same way produced them 37% of the time.
4. What our checks actually caught
Checks are only worth having if they find things. A sample of what ours found, and fixed:
- A method mismatch. The t-test page said it used Welch’s test; the server was running the pooled test. Both gave the same t-statistic on the page’s example, which is why it went unnoticed until a check compared degrees of freedom. The calculator now uses Welch by default.
- A wrong worked example. The same page quoted t ≈ 2.33 and p ≈ 0.049 for its example; the correct values are t = 2.40 and p = 0.046. Every number in our explanations is now recomputed rather than written by hand.
- A display that hid data. The Central Limit Theorem simulator cut its histogram at a fixed range, silently hiding 13% of the sample means in its most important scenario.
- A boundary rule. One tool treated a result as significant only when p < α; the standard convention, used by statsmodels, is p ≤ α.
- A wrong reference. Our first check of the Bradley–Terry standard errors disagreed with statsmodels — because the reference model was set up wrong, not the tool. Verifying the verifier matters too.
5. Behave sensibly at the edges
Real users type zero counts, identical values, mismatched lists and impossible combinations. A reliable tool explains what is wrong instead of returning a confident-looking number: our kappa calculator refuses when every label is identical (kappa is undefined), the Mann–Whitney and arena tools say when the data cannot support an answer, and the sequential A/B test warns when there are too few conversions for its approximation.
6. Be transparent and reproducible
A tool should say which method it used, show the formula or a worked example, and give the same answer every time. Our bootstrap tools use fixed random seeds so results are reproducible, and the methodology page lists how calculations are performed and validated.
A checklist for any free statistics tool
- Does it name the method? “95% CI” is not enough; which interval? Which t-test?
- Does it match a known answer? Try a textbook example or compare with R, SciPy or another tool you trust.
- Try the edges. 0 out of 20, 20 out of 20, two identical groups. Does it stay sensible?
- Exact or approximate? With small samples, approximations can mislead; good tools say which they use.
- Does it state assumptions — normality, independence, equal variances — and offer alternatives?
- Is it reproducible? The same input should give the same output, even for simulation-based methods.