The peeking problem
A standard significance test assumes you decide the sample size in advance and look at the result once. Real teams check dashboards daily and stop when the result looks good. Each extra look is another chance for noise to cross the threshold. We simulated 4,000 A/A tests (both variants converting at exactly 10%) and checked after every 200 visitors per variant, up to 100 looks: a standard z-test declared a “winner” in 37% of them. Looking only once at the planned end gave the expected 5%.
How always-valid p-values fix it
A sequential test accumulates evidence as a likelihood ratio that is guaranteed to stay small, with high probability, when there is no effect — no matter how many times you check. This calculator uses the mixture sequential probability ratio test (mSPRT) of Johari, Pekelis and Walsh, the method behind several commercial experimentation platforms. In the same 4,000 simulated A/A tests, peeking after every batch produced false positives in only 0.6–2.1% of them, below the 5% target.
Tuning: the smallest lift you care about
The “smallest lift” setting tunes the test’s sensitivity; it never affects validity. Set it near the effect you realistically hope to detect. Too small or too large mostly costs speed.
What it costs
Flexibility is not free. To detect a lift from 10% to 12%, a fixed-sample test needs 3,839 visitors per variant (see the A/B sample size calculator), and you must wait for all of them. In our simulation the sequential test detected that lift in every run, stopping after a median of 4,600 visitors per variant — slightly more on average, but you can stop early when the effect is large and you never have to fix the sample size in advance.
Common mistakes
- Using a standard calculator on a test you check daily. Its p-value is only valid for one pre-planned look.
- Stopping for a strong novelty effect. Run at least one full weekly cycle so weekday and weekend visitors are both represented.
- Changing the traffic split or the variant mid-test. That changes the experiment; start a new one.
- Running forever. Set a maximum duration; if the test has not concluded by then, the effect is probably smaller than you care about.
Related
Frequently asked questions
Not with a standard significance test: repeated checking inflates false positives (37% in our simulation of 100 looks). Use a sequential test with always-valid p-values, which stays valid however often you look.
A p-value from a sequential test that remains valid at any stopping time, so you can monitor continuously and stop as soon as it falls below your significance level.
On average it needs a somewhat larger sample to reach the same power, but it can stop early when the effect is large and does not require fixing the sample size in advance.
Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page.