What frequentist statistics is

The frequentist approach is the classical reading of probability: the probability of an event is the
share of times it occurs when the experiment is repeated many times over. It is an objectivist
statistics — it carries no prior beliefs and rests only on the observed data.

In A/B testing, the frequentist approach runs as a three-step procedure:

  1. Before the test: set α (the significance level, usually 0.05) and β (the acceptable Type II
    error, usually 0.2, i.e. 80% power), then compute the required sample size from the formula
  2. During the test: wait, without intervening, until the full sample is collected
  3. After the test: read the p-value and make the call

Key metrics

H₀ (null hypothesis):        CR(A) = CR(B), no difference
H₁ (alternative hypothesis): CR(B) ≠ CR(A)

α = 0.05 (Type I error — a false positive)
β = 0.20 (Type II error — a false negative)
Power = 1 − β = 0.80

p-value < α → reject H₀ → the result is "statistically significant"
Metric What it means
p-value The probability of this data given that there is no effect
α (alpha) The false-positive threshold (standard: 0.05)
Power The chance of detecting a real effect when one exists
Confidence interval The range containing the true difference with probability (1 − α)

The core discipline: no peeking

The first rule of frequentist testing is that there are no interim decisions. Look at the results
every day and stop the test the moment the gap looks attractive, and you inflate the Type I error:

Interim looks:          1     5     10     20
Real Type I error:    ~5%   ~14%   ~19%   ~25–30%
(at a nominal α = 0.05)

The fix is either strict discipline with no interim looks at all, or a move to sequential testing
with corrected boundaries.

Frequentist vs Bayesian: choosing between them

Both paradigms answer the same question, by different means:

Criterion Frequentist Bayesian
Interpretability for the business Harder (the p-value is unintuitive) Easier (“the probability that B is better”)
Early stopping Requires sequential testing Native
MAB / autopilot Not directly compatible Natively supported
Regulatory requirements The standard in pharma and finance Less widely accepted
Reproducibility High Depends on the prior

Tip: for most e-commerce experiments the Bayesian approach is the more practical one — it lets
you react sooner, it does not demand rigid peeking discipline, and it supports automatic traffic
allocation natively. Frequentist is preferable when you need formal reproducibility or have to
hand the result to external auditors.