Variation and control: the basic structure

Every A/B test has at least two groups:

  • Control (A) — the current version, unchanged. The baseline the effect is measured against.
  • Variation (treatment, B) — the version with the change. The goal is to measure whether that
    change improves the key metric.

Users are distributed between the groups at random and stay locked into their group for the whole
test (sticky assignment).

The single-change principle

The key requirement: one variation, one change. It is the only way to establish a causal link.

Hypothesis Correct variation Incorrect variation
“The image affects CR” A different photo only A different photo + copy + button colour
“The headline affects CTR” A different headline only A different headline + a rewritten description
“Widget position matters” A different position only Position + widget design

Breaking the principle does not make the test useless — you can still see the combined effect of a
bundle of changes. But you cannot identify what caused it.

A/B/n: several variations

In an A/B/n test one control is compared against several variations at the same time:

Control (A): the current recommendation algorithm
Variation B: collaborative filtering
Variation C: content-based
Variation D: blended strategy (50/50)

The upside is the time saved against testing sequentially. The downside is that every variation needs
its own quota of traffic: with 4 groups and 10K users needed per group, the test needs 40K in total.

Traffic allocation

By default traffic is split evenly — 50/50 for an A/B test, 25/25/25/25 for an A/B/C/D. An uneven
split (80/20 in favour of the control, for instance) is used when the variation carries real risk:
fewer users are exposed to a potentially negative experience.