Enter visitors and conversions for both variants to get conversion rates, lift, and whether the result is statistically significant.
The calculator runs a two-tailed z-test on two proportions. It pools the conversion rates of both variants, works out how much random variation you would expect from samples that size, and asks how far the observed gap sits outside that noise. The output is confidence: the probability that a difference this large would not have appeared by chance alone.
| Baseline rate | To detect +10% relative |
|---|---|
| 1% | ~155,000 per variant |
| 2% | ~77,000 per variant |
| 5% | ~30,000 per variant |
| 10% | ~14,000 per variant |
| 20% | ~6,000 per variant |
An inconclusive test is not a failed test – it is a result. It says the difference between the two variants, if any exists, is smaller than what you can detect with the traffic you have. That is useful: it tells you to stop spending attention on this change and go find a bigger lever.
Peeking.
Peeking. You watch the dashboard daily, and the moment confidence crosses 95 percent you stop the test and declare a winner. Confidence wanders as data accumulates – it will cross 95 percent by accident on a coin flip given enough looks. Stopping the first time it does turns a 5 percent error rate into something closer to 30 percent.
Decide the sample size before the test starts, run to it, then read the result once. If you must monitor, monitor for disasters, not for wins.