What a Type I error is
In A/B testing the null hypothesis (H₀) states that variations A and B are identical. A Type I error
is rejecting H₀ when it is in fact true. Put simply: the test “found a winner” that does not exist.
The parameter α controls the acceptable probability of that error. At α = 0.05, in a test with no
real effect, one experiment in twenty will return a false positive purely by chance.
Practical consequences
For a team running 40 tests a year at α = 0.05:
Expected number of false "wins" = 40 × 0.05 = 2 per year
That means two rollouts of changes that do not actually improve the product. With a small MDE
(minimum detectable effect) and low traffic, the real share of false positives can be higher still.
The main amplifiers of Type I error
Peeking. Check the results daily and stop at the first p < 0.05, and the real probability of a
Type I error across 20 interim checks exceeds 30% — against a stated α of 0.05. This is the single
most common mistake in A/B testing.
Multiple metrics without a correction. Evaluate a test on 10 metrics and the probability of a
chance significant result on at least one is already around 40%. The fix is to pick one primary
metric in advance and apply a Bonferroni correction or FDR control to the rest.
Sample Ratio Mismatch (SRM). When the actual group ratio differs from the planned one, the
results are distorted and the significance can be spurious.
Tip: the Bayesian approach does not use a p-value or an α threshold directly — it reports the
probability to be best instead. That does not remove Type I error entirely, but it makes it far
less dependent on how many times you looked at the data.
Type I versus Type II error
| Type I error (α) | Type II error (β) | |
|---|---|---|
| What happens | We roll out a bad change | We reject a good change |
| Controlled by | The α level | Test power (1 − β) |
| Reduced by | A smaller α, sequential testing | More traffic, a smaller MDE |
| Trade-off | A stricter α raises β | A stricter β raises α |