How MVT differs from an A/B test

In an A/B test a single element changes — a headline, an image or a recommendation algorithm — and
two versions are compared. In MVT several independent elements change at the same time, and every
combination of them is tested.

An example on a product page, with two factors:
– Headline of the recommendation block: A (current) / B (new)
– Position of the block: above the fold / below the fold

That gives four cells: A-above, A-below, B-above, B-below, plus the control. It answers not only
“which version is better” but also “does the headline effect depend on where the block sits”.

Full factorial vs fractional factorial

Full factorial — every combination is tested. More reliable, but it demands the most traffic.

Fractional factorial — only part of the combinations is tested, following a designed plan. It
estimates the main effects of each factor on less traffic, but gives up the ability to measure some
of the interactions.

Approach Cells (2 factors × 3 versions) Cells (3 factors × 3 versions)
Full factorial 9 27
Fractional factorial 5–6 9–12

Important: every cell needs a full sample of its own. If an ordinary A/B test needs 10,000
users per variation, an MVT with 9 cells needs 90,000. Without enough traffic an MVT will stretch
over months or never reach a significant result at all.

When to run MVT and when to run sequential A/B tests

The rule is straightforward:

  • A hypothesis about interaction between factors → MVT. “The effect of the new headline depends
    on what the image looks like” is an interaction, and an A/B test will never surface it.
  • Independent hypotheses → sequential A/B tests. Faster, simpler, less traffic.
  • Traffic under 500K MUV a month → almost always sequential A/B. MVT will simply take too long.

Common mistakes

  • Too many factors. 4 factors × 3 versions = 81 cells. At 1M MUV each cell gets roughly 12K users
    a month — not enough for most CR questions.
  • Analysis without a multiple-comparisons correction. With 27 cells, the chance of finding a
    “significant” result purely from noise is very high.
  • Ignoring interactions. Skip the interaction effects and you may roll out a winning cell that
    only works in that one specific combination.