Why randomization is the basis of a trustworthy test

An A/B test measures a causal link: this change is what lifted conversion. That claim only holds if
groups A and B are comparable — that is, if they differ only in the change under test and not in the
make-up of the audience.

Randomization is what delivers that. With a correct random split, the groups are statistically
identical on every characteristic — new versus returning users, devices, traffic channels, time of day.

Randomization methods

Hash function (the standard for web and mobile):

# Deterministic randomization through a hash
import hashlib

def get_variant(user_id: str, test_salt: str, num_variants: int) -> int:
    key = f"{user_id}:{test_salt}"
    hash_value = int(hashlib.md5(key.encode()).hexdigest(), 16)
    return hash_value % num_variants

# user_id="u12345", test_salt="test_homepage_hero", num_variants=2
# -> 0 (control) or 1 (variation)

The test salt (test_salt) is unique to each experiment. That is what keeps splits independent: one
user can sit in the control of one test and in the variation of another.

Units of randomization

Unit When to use it Risks
User (userId) The standard for signed-in audiences Needs separate logic for anonymous users
Cookie For anonymous users Unstable when the browser changes
Session Tests that do not carry behaviour across sessions One person sees both variations
Page / request Backend load tests Unsuitable for UX metrics

Important: the unit of randomization has to match the unit of measurement. If the metric is
“conversion per user”, randomize per user — otherwise you introduce a statistical bias.

Typical breaches of randomization

  • Splitting by time — week A versus week B. Imports seasonality and events as confounders.
  • Splitting by geography — city A versus city B. Differences come from the audience, not the test.
  • Randomizing per session while measuring per user — the user sees both variations and the effect washes out.
  • The same salt across every test — the splits correlate, and parallel experiments interfere with each other.

Verifying randomization: A/A tests and SRM

Before the first A/B test it is worth running an A/A test — it demonstrates that randomization
produces no significant difference between two identical groups.

While tests run, watch Sample Ratio Mismatch (SRM): if you expected a 50/50 split and got 53/47,
either randomization is broken or traffic is being filtered unevenly.