What frequentist statistics is
The frequentist approach is the classical reading of probability: the probability of an event is the
share of times it occurs when the experiment is repeated many times over. It is an objectivist
statistics — it carries no prior beliefs and rests only on the observed data.
In A/B testing, the frequentist approach runs as a three-step procedure:
- Before the test: set α (the significance level, usually 0.05) and β (the acceptable Type II
error, usually 0.2, i.e. 80% power), then compute the required sample size from the formula - During the test: wait, without intervening, until the full sample is collected
- After the test: read the p-value and make the call
Key metrics
H₀ (null hypothesis): CR(A) = CR(B), no difference
H₁ (alternative hypothesis): CR(B) ≠ CR(A)
α = 0.05 (Type I error — a false positive)
β = 0.20 (Type II error — a false negative)
Power = 1 − β = 0.80
p-value < α → reject H₀ → the result is "statistically significant"
| Metric | What it means |
|---|---|
| p-value | The probability of this data given that there is no effect |
| α (alpha) | The false-positive threshold (standard: 0.05) |
| Power | The chance of detecting a real effect when one exists |
| Confidence interval | The range containing the true difference with probability (1 − α) |
The core discipline: no peeking
The first rule of frequentist testing is that there are no interim decisions. Look at the results
every day and stop the test the moment the gap looks attractive, and you inflate the Type I error:
Interim looks: 1 5 10 20
Real Type I error: ~5% ~14% ~19% ~25–30%
(at a nominal α = 0.05)
The fix is either strict discipline with no interim looks at all, or a move to sequential testing
with corrected boundaries.
Frequentist vs Bayesian: choosing between them
Both paradigms answer the same question, by different means:
| Criterion | Frequentist | Bayesian |
|---|---|---|
| Interpretability for the business | Harder (the p-value is unintuitive) | Easier (“the probability that B is better”) |
| Early stopping | Requires sequential testing | Native |
| MAB / autopilot | Not directly compatible | Natively supported |
| Regulatory requirements | The standard in pharma and finance | Less widely accepted |
| Reproducibility | High | Depends on the prior |
Tip: for most e-commerce experiments the Bayesian approach is the more practical one — it lets
you react sooner, it does not demand rigid peeking discipline, and it supports automatic traffic
allocation natively. Frequentist is preferable when you need formal reproducibility or have to
hand the result to external auditors.