What a p-value is and how to read it

A p-value is the probability of obtaining the observed result, or a more extreme one, assuming the
null hypothesis is true. In an A/B test the null hypothesis is that there is no difference between
variation A and variation B.

Formally: p = 0.03 means that if no difference existed, we would see an effect this large or larger
by chance in 3% of repeated experiments.

What a p-value does not mean:

  • 96% probability that variation B is better — no
  • The null hypothesis is false with 96% probability — no
  • The effect is practically meaningful — no, only statistically detectable

Three things that break the interpretation

Peeking — checking interim results with the option to stop. Every additional check is
effectively another test on the same data, which inflates the real error rate. Twenty checks at 5%
risk each puts the chance of a spurious p < 0.05 at roughly 64%.

A small sample — with insufficient data a high p-value does not mean no effect, only not enough
data to detect one. That is a type II error.

Multiple hypotheses — testing 20 metrics at once at p < 0.05 means one of them will look
significant by chance on average. Bonferroni and other multiple-comparison corrections reduce that
risk.

Frequentist versus Bayesian: the real difference

Property Frequentist (p-value) Bayesian (probability to be best)
What it measures Probability of the data under H₀ Probability that B beats A
Interpretation Technical, frequently misread Direct
Early stopping Invalid Valid
Sample size Must be fixed in advance Flexible

Important: a p-value is not a probability of winning. They are different quantities. The
Bayesian approach reports the probability of winning directly, which is both easier to read and
safer when results are being watched during the test.