Why randomization is the basis of a trustworthy test
An A/B test measures a causal link: this change is what lifted conversion. That claim only holds if
groups A and B are comparable — that is, if they differ only in the change under test and not in the
make-up of the audience.
Randomization is what delivers that. With a correct random split, the groups are statistically
identical on every characteristic — new versus returning users, devices, traffic channels, time of day.
Randomization methods
Hash function (the standard for web and mobile):
# Deterministic randomization through a hash
import hashlib
def get_variant(user_id: str, test_salt: str, num_variants: int) -> int:
key = f"{user_id}:{test_salt}"
hash_value = int(hashlib.md5(key.encode()).hexdigest(), 16)
return hash_value % num_variants
# user_id="u12345", test_salt="test_homepage_hero", num_variants=2
# -> 0 (control) or 1 (variation)
The test salt (test_salt) is unique to each experiment. That is what keeps splits independent: one
user can sit in the control of one test and in the variation of another.
Units of randomization
| Unit | When to use it | Risks |
|---|---|---|
| User (userId) | The standard for signed-in audiences | Needs separate logic for anonymous users |
| Cookie | For anonymous users | Unstable when the browser changes |
| Session | Tests that do not carry behaviour across sessions | One person sees both variations |
| Page / request | Backend load tests | Unsuitable for UX metrics |
Important: the unit of randomization has to match the unit of measurement. If the metric is
“conversion per user”, randomize per user — otherwise you introduce a statistical bias.
Typical breaches of randomization
- Splitting by time — week A versus week B. Imports seasonality and events as confounders.
- Splitting by geography — city A versus city B. Differences come from the audience, not the test.
- Randomizing per session while measuring per user — the user sees both variations and the effect washes out.
- The same salt across every test — the splits correlate, and parallel experiments interfere with each other.
Verifying randomization: A/A tests and SRM
Before the first A/B test it is worth running an A/A test — it demonstrates that randomization
produces no significant difference between two identical groups.
While tests run, watch Sample Ratio Mismatch (SRM): if you expected a 50/50 split and got 53/47,
either randomization is broken or traffic is being filtered unevenly.