What power is

Power (1 − beta) is the probability that a test correctly detects a real effect. Its complement,
beta, is the probability of a type II error: missing a genuine improvement.

Reality:              B is better         B is not better
Test says "better":   Correct (power)     Type I error (alpha)
Test says "no":       Type II error (beta) Correct

The industry standard is power at 80% or above (beta at 20% or below) with alpha at 0.05. Out of 100
tests where B really is better, 80 are correctly identified and 20 are missed.

What drives power

Sample size. The main lever. More users in the test means less random noise and a better chance
of seeing a real effect.

Effect size (MDE). A large effect (+20% CR) is easier to detect at the same sample than a small
one (+3%). That is why a realistic estimate of the minimum meaningful effect belongs in the plan.

The baseline metric. At a 5% conversion rate the variance is lower than at 0.5%, so a smaller
sample achieves the same power.

Calculating the sample you need

For a standard test at alpha 0.05 and 80% power the formula needs:

  • the baseline conversion rate
  • the MDE, the minimum detectable effect, say 10%

Approximate relationships:

Baseline CR MDE Traffic per variation
2% 15% ~6,500
2% 10% ~14,700
1% 15% ~13,000
1% 10% ~29,400

Important: these numbers are per variation. With two variations the total is double.

Power and Bayesian statistics

Power in the classical sense is a frequentist instrument. Bayesian statistics do not require power to
be fixed in advance: evidence accumulates in favour of each hypothesis instead. Probability to be
best rises as data arrives, with no peeking risk and no formula-driven sample plan.

For teams with limited traffic, the Bayesian route allows a grounded decision earlier — when the
data is still short of 80% frequentist power but already sufficient for a confident probabilistic
read.