What SRM is and why it is dangerous
Sample Ratio Mismatch (SRM) is the situation where users did not split between the variations of a
test in the configured ratio. You set a 50/50 split, expect 50,000 users in each group, and end up
with 48,000 in the control and 52,000 in the variation.
The danger is not the deviation itself but the cause behind it: it is a symptom of a systematic error
in the test’s logic. When the groups are formed incorrectly they cannot be treated as comparable, so
the difference in metrics cannot be attributed to the change.
Important: SRM makes a test’s results unreliable in principle. A winner declared in a test with
SRM may be an artefact of the bug rather than a real improvement.
How to detect SRM: the chi-squared test
The standard method is the chi-squared criterion:
H0: the observed distribution matches the expected one
Observed: [48,000, 52,000]
Expected: [50,000, 50,000]
chi2 = (48000-50000)^2/50000 + (52000-50000)^2/50000
chi2 = 80 + 80 = 160
p-value << 0.001 -> SRM confirmed
The rule: on any test, check SRM first and analyse the metrics second.
Causes of SRM
| Cause | How it shows up | Fix |
|---|---|---|
| Bot traffic | Bots land in only one variation | Filter bots at the data layer |
| Asymmetric load speed | The slower variation bounces more | Optimise performance |
| A bug in the splitting logic | Cookies get overwritten, users “jump” groups | Check sticky assignment |
| Caching | The CDN serves one variation to part of the audience | Exclude cacheable requests from the split |
| Duplicated events | One session is counted twice | Deduplicate in the pipeline |
SRM across different test types
UI A/B tests: especially exposed to SRM through load-speed differences — the slower variation
produces more bounces.
A/B tests of recommendation algorithms: less exposed to speed-driven SRM, but they can suffer
from an uneven split of new versus returning users at the start of the test.
MAB (multi-armed bandit): dynamic traffic allocation changes the proportions by design — that is
normal. The standard SRM check does not apply to MAB.
How to prevent SRM
- An A/A test before the first experiment — confirm the system forms groups correctly under neutral conditions
- Monitor group sizes from the first hours of a test — an early deviation is easier to find before large volumes of data pile up
- Sticky assignment — once a user lands in a group they see the same variation on every subsequent visit
- Bot filtering at the point of data collection, not after the fact