This calculator tells you how many visitors each variant of an A/B test needs before the result is trustworthy. You give it your current (baseline) conversion rate, the new rate you hope to detect, your significance level (usually 5%), and your power (usually 80%). Smaller improvements and higher confidence require more traffic — halving the effect you want to detect roughly quadruples the sample. Deciding the size in advance, and not stopping early, is what keeps an A/B test honest.

When to use it

Run this before you launch an A/B test — a landing-page variant, an email subject line, a checkout change. It converts your business question (“is the new version better?”) into the traffic you must collect before the answer is reliable. Skipping this step is why so many tests are called early on noise.

The inputs explained

  • Baseline rate — your current conversion rate for the control.
  • Expected new rate — the rate for the variant; the gap between the two is your minimum detectable effect (MDE). Set it to the smallest lift that would actually change your decision, not a hopeful guess.
  • Significance level (α) — the false-positive rate you tolerate; 0.05 is standard.
  • Power — the chance of detecting a real effect; 80% is standard, 90%+ for high-stakes tests.

Worked example

A checkout converts at 10%. You want to detect an improvement to 15% at α = 0.05 and 80% power. The calculator returns about 683 visitors per variant (roughly 1,366 total). If instead you only cared about a jump to 11%, the requirement balloons into the tens of thousands — small effects are expensive to prove.

The formula

For two proportions, the sample size per group is

n = (z1−α/2 + z1−β)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁ − p₂)²

where z values are standard-normal quantiles. This is the normal approximation used by all standard A/B calculators; round up to a whole number. See the methodology page.

Common mistakes

  • Peeking — checking daily and stopping at the first “significant” moment inflates false positives far above 5%. Run to the planned size.
  • An unrealistically large MDE — assuming a huge lift gives a comfortably small sample but a test that misses the real, smaller effect.
  • Ignoring the business question — a statistically significant 0.1% lift may not be worth shipping; weigh it against cost.
  • One variant only — remember you need this many per variant.

Related

Two-Proportion TestSample Size (means)Survey Sample SizeType I & II errors