What a confidence interval is
Any number produced by an experiment is an estimate, not the truth. A measured 2% conversion
gain carries error: the real effect could be somewhat larger or smaller. A confidence interval
formalises that error.
95% CI for the CR gain: [+0.8%, +3.4%]
→ the true effect most plausibly lies between +0.8% and +3.4%
→ the lower bound is above zero → the difference is significant
95% is the standard confidence level in A/B testing. It corresponds to the p < 0.05 threshold in
frequentist statistics.
Why an interval beats a p-value
A p-value answers one question: is there a difference? A confidence interval answers two: is there a
difference, and how large is it?
| Situation | p-value | CI | Conclusion |
|---|---|---|---|
| Large effect, large sample | p < 0.001 | [+5%, +9%] | Significant and practically important |
| Tiny effect, enormous sample | p = 0.03 | [+0.01%, +0.3%] | Significant, practically irrelevant |
| No effect | p = 0.52 | [−2%, +3%] | Not significant, effect undetermined |
A statistically significant but practically meaningless result on a huge sample is a common trap in
e-commerce. The p-value announces a winner; the interval reveals that the real gain is too small to
pay for the rollout.
Reading an interval in an A/B test
The interval excludes zero → the difference is statistically significant. B differs from A.
The interval includes zero → there is no basis to claim B is better. Either keep accumulating
sample or stop the test.
Tip: treat the lower bound as the pessimistic scenario. If even the worst case pays back the
cost of the change, ship it.
Width and sample size
Interval width is inversely proportional to the square root of the number of observations:
- At 1,000 users per group the interval may be ±5%
- At 10,000 — ±1.6%
- At 100,000 — ±0.5%
An insufficient sample means a wide interval means no clear conclusion. That is why computing the
sample size before the test is mandatory rather than optional.
Common misreadings
- The interval contains the true value with 95% probability — imprecise. Correctly: across
repeated experiments, 95% of such intervals contain the true parameter. - Ignoring the width — reading only the point estimate (+2.1%) while the interval
[−0.5%, +4.7%] still includes negative outcomes. - Stopping at the first attractive interval — under peeking the interval narrows temporarily and
may exclude zero, but that is an artefact of a small sample.