What a Type II error is

In A/B testing a Type II error is accepting the null hypothesis — H₀, “the variations are identical” —
when variation B is actually better. The test says “no difference” while a difference exists.

It is denoted β and it is the complement of power: power = 1 − β. At the standard 80% power:

β = 0.20 → one test in five with a real effect will report "no result"

Why it happens

The main cause is an insufficient sample. When traffic is thin or the test is stopped too early, the
statistics never get the chance to see the real effect. That matters most for small effects: a 3% CR
improvement — modest but valuable — needs several times more data than a 15% one.

The second cause is a poorly chosen metric. Conversion rate (CR) is a noisy metric with high variance.
A more sensitive alternative is RPV (revenue per visitor): it accounts for conversion and order value
at the same time, which lowers β at the same sample size.

How to reduce Type II error

Enlarge the sample. The most direct route: collect data for longer, or pick higher-traffic pages
to test on.

Choose a more sensitive metric. RPV instead of CR, AOV for assortment tests, attributed revenue
for recommendations.

Lower the MDE. Sometimes the question is not “how do we enlarge the sample” but “what is the
smallest effect we actually care about”. If a 2% conversion gain is too small to justify the size of
the test, the problem may be framed wrongly.

Use CUPED (controlled experiment using pre-experiment data). The method reduces the variance of a
metric by adjusting for pre-experiment covariates, raising power at the same sample size.

Important: lowering β always trades against α. Raise power by loosening α and the Type I error
grows. The industry standard is α = 0.05 and β = 0.20 (80% power). Depart from it deliberately, not
by accident.

Comparing the two error types

Parameter Type I (α) Type II (β)
What went wrong We shipped a neutral change We did not ship a good change
What it affects The product (neutral changes) Growth (missed improvements)
Controlled by The significance threshold α Test power = 1 − β
Reduced by Sequential testing, a smaller α A larger N, a better metric, CUPED