How MVT differs from an A/B test
In an A/B test a single element changes — a headline, an image or a recommendation algorithm — and
two versions are compared. In MVT several independent elements change at the same time, and every
combination of them is tested.
An example on a product page, with two factors:
– Headline of the recommendation block: A (current) / B (new)
– Position of the block: above the fold / below the fold
That gives four cells: A-above, A-below, B-above, B-below, plus the control. It answers not only
“which version is better” but also “does the headline effect depend on where the block sits”.
Full factorial vs fractional factorial
Full factorial — every combination is tested. More reliable, but it demands the most traffic.
Fractional factorial — only part of the combinations is tested, following a designed plan. It
estimates the main effects of each factor on less traffic, but gives up the ability to measure some
of the interactions.
| Approach | Cells (2 factors × 3 versions) | Cells (3 factors × 3 versions) |
|---|---|---|
| Full factorial | 9 | 27 |
| Fractional factorial | 5–6 | 9–12 |
Important: every cell needs a full sample of its own. If an ordinary A/B test needs 10,000
users per variation, an MVT with 9 cells needs 90,000. Without enough traffic an MVT will stretch
over months or never reach a significant result at all.
When to run MVT and when to run sequential A/B tests
The rule is straightforward:
- A hypothesis about interaction between factors → MVT. “The effect of the new headline depends
on what the image looks like” is an interaction, and an A/B test will never surface it. - Independent hypotheses → sequential A/B tests. Faster, simpler, less traffic.
- Traffic under 500K MUV a month → almost always sequential A/B. MVT will simply take too long.
Common mistakes
- Too many factors. 4 factors × 3 versions = 81 cells. At 1M MUV each cell gets roughly 12K users
a month — not enough for most CR questions. - Analysis without a multiple-comparisons correction. With 27 cells, the chance of finding a
“significant” result purely from noise is very high. - Ignoring interactions. Skip the interaction effects and you may roll out a winning cell that
only works in that one specific combination.