What power is
Power (1 − beta) is the probability that a test correctly detects a real effect. Its complement,
beta, is the probability of a type II error: missing a genuine improvement.
Reality: B is better B is not better
Test says "better": Correct (power) Type I error (alpha)
Test says "no": Type II error (beta) Correct
The industry standard is power at 80% or above (beta at 20% or below) with alpha at 0.05. Out of 100
tests where B really is better, 80 are correctly identified and 20 are missed.
What drives power
Sample size. The main lever. More users in the test means less random noise and a better chance
of seeing a real effect.
Effect size (MDE). A large effect (+20% CR) is easier to detect at the same sample than a small
one (+3%). That is why a realistic estimate of the minimum meaningful effect belongs in the plan.
The baseline metric. At a 5% conversion rate the variance is lower than at 0.5%, so a smaller
sample achieves the same power.
Calculating the sample you need
For a standard test at alpha 0.05 and 80% power the formula needs:
- the baseline conversion rate
- the MDE, the minimum detectable effect, say 10%
Approximate relationships:
| Baseline CR | MDE | Traffic per variation |
|---|---|---|
| 2% | 15% | ~6,500 |
| 2% | 10% | ~14,700 |
| 1% | 15% | ~13,000 |
| 1% | 10% | ~29,400 |
Important: these numbers are per variation. With two variations the total is double.
Power and Bayesian statistics
Power in the classical sense is a frequentist instrument. Bayesian statistics do not require power to
be fixed in advance: evidence accumulates in favour of each hypothesis instead. Probability to be
best rises as data arrives, with no peeking risk and no formula-driven sample plan.
For teams with limited traffic, the Bayesian route allows a grounded decision earlier — when the
data is still short of 80% frequentist power but already sufficient for a confident probabilistic
read.