The peeking problem — why you cannot simply look

A classical frequentist A/B test requires the sample size to be fixed in advance and then waited
for. Watching interim results and stopping once it looks significant is peeking. Here is why that
matters.

At a threshold of p < 0.05 the false positive rate is fixed at 5% — for one check. Repeated
checks accumulate:

1 check   → false positives ~5%
5 checks  → false positives ~19%
10 checks → false positives ~30%
20 checks → false positives ~54%

A team that opens the dashboard daily and stops on significance is effectively deciding from noise.

How sequential testing works

Sequential testing makes look whenever you like statistically legitimate in two ways.

1. An alpha spending function

The total type I error budget of 0.05 is spent according to a defined function across the interim
checks. Each check uses a stricter threshold than a single test would, and the total spend never
exceeds 0.05.

2. E-values and always-valid inference

An e-value is an accumulating measure of evidence:

e_t = e_{t-1} × LR_t  (the product of likelihood ratios at each step)

The test can stop when e_t ≥ 1/alpha. The key property is that validity holds under any
stopping strategy — including a discretionary one.

Sequential versus Bayesian early stopping

Property Sequential (e-values) Bayesian
Type I error control Strict frequentist Probabilistic
Interpretation There is enough evidence The probability B wins is X%
Clarity for the business Medium High
Formal rigour Very high Depends on the prior

For most product teams the Bayesian approach is clearer and sufficient. Sequential testing with
e-values is the instrument for data science teams with strict formal requirements.

When to use it

  • High-stakes tests where mistakes are expensive (homepage, checkout)
  • Tests under unstable traffic (seasonality, promotions)
  • Teams that cannot wait for a fixed end date
  • Product organisations with a mature experimentation culture