Free tool

A/B/n Test Calculator

Check the statistical significance of your experiment — up to 4 variants, two methods: frequentist (p-value, CI) and Bayesian (probability to win).

  1. 01
    Enter your data Visitors and conversions for every variant
  2. 02
    Pick a method Frequentist is stricter, Bayesian is faster
  3. 03
    Read the result Green = significant, amber = not enough data yet

Test results

Analysis settings

Variant data A is the control

2 of 4
Variant Visitors Conversions CR, %

Up to 4 variants (A, B, C, D). With 3 or more variants, account for the multiple comparisons correction.

Results compared with control A

Test planning

Sample size before the test starts

Required sample size

Methodology

Bayesian vs Frequentist

Bayesian · Beta-Binomial

Faster operational decisions

  • Answers: “how likely is it that B beats A?”
  • Decision threshold — P(B > A) ≥ 95%
  • Can be evaluated at any point while data accumulates
  • Prior — uniform Beta(1,1), neutral
  • Expected uplift is estimated from 20,000 simulations

Better for: fast e-commerce decisions, small and mid-sized samples.

Frequentist · z-test of proportions

Strict hypothesis testing

  • Returns a p-value: the chance of seeing this difference under H₀
  • Confidence interval for the absolute effect
  • Sample size must be fixed before the test starts
  • No early stopping — it inflates the false positive rate
  • Neutral to any prior knowledge about the conversion rate

Better for: large samples, strict financial or regulatory decisions.

FAQ

Frequently asked questions

The p-value is the probability of observing a difference at least as large as the one you see, assuming there is no real effect between the variants (the null hypothesis). A p-value below 0.05 is conventionally called statistically significant: the chance of a purely random result is under 5%. Important: the p-value says nothing about the size of the effect or its practical significance.

It means that with the data collected so far you cannot claim with confidence that the difference is not down to chance. It does not mean the variant is worse or that the test failed — only that it is too early to conclude. Collect more data or revisit the minimum effect you care about.

The minimum sample size depends on the baseline conversion rate, the minimum detectable effect (MDE) and the significance level. For typical e-commerce (CR ~2–5%, MDE ~10–15%) you need 5,000 to 20,000 visitors per variant. Use the “Sample size” block above — it calculates the number for you.

The frequentist approach assumes a fixed sample size decided in advance. If you keep peeking at the results and stop as soon as p < 0.05, the real type I error rate ends up higher than the α you declared. With 5 interim looks the false positive risk grows from 5% to roughly 22%. The Bayesian approach lets you make a call at any point.

A 95% CI means that if you repeated the test many times, in 95% of those repetitions the interval would cover the true absolute effect. If the interval does not cross zero, the effect is statistically significant. The narrower the interval, the more precise the estimate — which usually takes a larger sample.

When you compare A vs B, A vs C and A vs D at the same time, the chance of at least one false positive grows. Across three tests at α = 5% the combined risk of a spurious result reaches 14%. The standard correction is Bonferroni: divide α by the number of comparisons. For three variants use α / 2 = 2.5%. The calculator shows a warning with 3 or more variants.

MDE (Minimum Detectable Effect) is the smallest relative lift in conversion that is practically meaningful for the business. For example, if it matters to you to detect a CR lift from 6.5% to 7.15%, the MDE is 10%. The smaller the MDE, the larger the sample. E-commerce teams usually set an MDE of 5–15%: anything smaller is hard to detect without a very long test.

The tool implements a two-sample z-test of proportions and a Bayesian Beta-Binomial model. It only fits binary metrics (conversion, click, purchase). Not supported: continuous metrics (average order value), CUPED and other variance reduction methods. With a baseline conversion rate below 1%, or samples under 100 per variant, the normal approximation is unreliable.

Gravity Field

Looking for an A/B testing platform for e-commerce?

Gravity Field builds the A/B engine into your personalization stack: audience segmentation, automatic traffic allocation and real-time result analytics.

Request a demo