What Makes a Free Online Statistics Tool Reliable and Accurate?

By Dackohn · 2026-10-06

A statistics tool is reliable when it uses the right method for the situation, matches established reference software on known examples, behaves sensibly at the edges, and tells you exactly what it computed. Arithmetic is rarely the problem; choosing the method is. This post explains how we check the calculators on this site, with real numbers — including the mistakes the checks caught — and ends with a checklist you can apply to any free tool.

1. The method matters more than the arithmetic

Two calculators can both compute flawlessly and still disagree, because they use different methods. Two examples:

Confidence intervals for a proportion. The textbook “Wald” interval, p ± 1.96√(p(1−p)/n), is still the default in many tools. We computed its exact coverage — how often a 95% interval actually contains the true value — by summing binomial probabilities:

Sample size, true proportionWald coverageWilson coverage
n = 20, p = 0.5095.9%95.9%
n = 20, p = 0.9087.6%95.7%
n = 20, p = 0.9563.9%92.5%
n = 100, p = 0.9886.6%94.9%

Near 0% or 100% the “95%” Wald interval can be wrong a third of the time; at 100 successes out of 100 it collapses to zero width. That is why our proportion intervals use Wilson.

Comparing two means. Student’s classic t-test assumes both groups have the same variance; Welch’s version does not. In one test case — a group of 8 with a small spread against a group of 25 with a large one — Welch gave p = 0.023 and the pooled test p = 0.130: opposite conclusions from the same data. Welch is the safer default, costing almost nothing when the variances happen to be equal.

2. Check every result against reference software

We test each calculation against established scientific libraries — SciPy, statsmodels, scikit-learn and specialist packages — on hundreds of random inputs, not just one example. The browser-based tools are tested by running their actual JavaScript in a separate engine and comparing outputs. Typical worst-case differences:

The point is not the decimals but the breadth: a tool that matches on random inputs, edge cases included, is far more trustworthy than one that matches a single textbook example.

3. Test the claims, not just the formulas

Some promises can only be checked by simulation. Our sample-size calculator promises 80% power; in 20,000 simulated studies at the recommended size, the test detected the effect 79.7–80.1% of the time. The sequential A/B test promises that checking results every day will not inflate false positives; in 4,000 simulated tests with no real difference, peeking 100 times produced false positives 0.6–2.1% of the time — while a standard test checked the same way produced them 37% of the time.

4. What our checks actually caught

Checks are only worth having if they find things. A sample of what ours found, and fixed:

5. Behave sensibly at the edges

Real users type zero counts, identical values, mismatched lists and impossible combinations. A reliable tool explains what is wrong instead of returning a confident-looking number: our kappa calculator refuses when every label is identical (kappa is undefined), the Mann–Whitney and arena tools say when the data cannot support an answer, and the sequential A/B test warns when there are too few conversions for its approximation.

6. Be transparent and reproducible

A tool should say which method it used, show the formula or a worked example, and give the same answer every time. Our bootstrap tools use fixed random seeds so results are reproducible, and the methodology page lists how calculations are performed and validated.

A checklist for any free statistics tool

  1. Does it name the method? “95% CI” is not enough; which interval? Which t-test?
  2. Does it match a known answer? Try a textbook example or compare with R, SciPy or another tool you trust.
  3. Try the edges. 0 out of 20, 20 out of 20, two identical groups. Does it stay sensible?
  4. Exact or approximate? With small samples, approximations can mislead; good tools say which they use.
  5. Does it state assumptions — normality, independence, equal variances — and offer alternatives?
  6. Is it reproducible? The same input should give the same output, even for simulation-based methods.