Why peeking is the most common mistake in A/B testing

The test is running, day three. You open the dashboard and see variation B ahead of A at p = 0.03.
Looks like a winner. Can you stop?

No. That is exactly what peeking is.

How a p-value behaves over time

A p-value does not descend monotonically toward its correct final value. It fluctuates — it can
drop below 0.05 on day three, rise to 0.15 on day five, fall to 0.04 on day eight. Checking every
day and stopping at the first significant reading means taking a random point from that walk.

Day 3:  p = 0.03  ← "winner!" (still noise)
Day 5:  p = 0.14  ← what you would have seen had you waited
Day 8:  p = 0.04  ← "significant" again
Day 14: p = 0.22  ← the final result: no effect

Across 20 daily checks the probability of seeing p < 0.05 at least once is around 64%, with a true
effect of zero.

How many false wins peeking creates

Number of interim checks Real type I error rate
1 (no peeking) 5%
5 ~22%
10 ~40%
20 ~64%

Under aggressive peeking, most significant results are noise.

Working with interim data correctly

The Bayesian approach is the most practical answer. Probability to be best is interpretable at
any moment and does not break when results are watched along the way.

Sequential testing is the frequentist route: apply alpha spending, which defines in advance how
the error budget is consumed at each check.

Tip: decide before launch when the test will stop. Fix the planned duration or the sample size.
The rule of only looking at the end is the simplest and, where the discipline exists, the most
effective.