Testing math

A/B Test Significance Calculator

Before you ship the winner, check whether your test found a real difference or just measured noise.

Significance uses a two-proportion z-test: z = (rateB − rateA) ÷ standard error of the difference. Suppose variant A converts 150 of 5,000 visitors (3.0%) and variant B converts 180 of 5,000 (3.6%): the lift is 20%, the z-score is 1.68, and confidence comes out near 91% — promising, but short of the conventional 95% bar, so more traffic is needed before declaring a winner. This calculator uses the normal approximation, which is reliable once each variant has at least a few dozen conversions. Enter both variants' visitors and conversions to get lift, z-score and confidence.

A/B Test Significance Calculator — your numbers

Lift (B vs A)

20.0%

Z-score

1.68

Confidence

90.7%

Estimate only. Results reflect exactly the numbers you enter — verify against your own accounting before making pricing decisions.

Embed this calculator on your site — free →

Worked examples

Promising but not proven

Visitors — variant A 5000
Conversions — variant A 150
Visitors — variant B 5000
Conversions — variant B 180
Lift (B vs A) 20.0%
Z-score 1.68
Confidence 90.7%

A 20% lift at ~91% confidence — keep the test running to reach 95%.

Clear winner

Visitors — variant A 8000
Conversions — variant A 240
Visitors — variant B 8000
Conversions — variant B 320
Lift (B vs A) 33.3%
Z-score 3.44
Confidence 99.9%

A 33% lift with z ≈ 3.4 — beyond 99.9% confidence; ship variant B.

Frequently asked questions

What confidence level should I require before calling a winner?

95% is the standard bar — it means a difference this large would appear by chance in fewer than 1 in 20 no-difference tests. For low-risk changes like button copy, some teams accept 90%; for pricing or checkout changes where a wrong call is expensive, hold out for 95% or higher plus a sensible minimum run time of two full business cycles.

Why should I not peek at results and stop early?

Checking repeatedly and stopping the moment confidence crosses 95% inflates your false-positive rate badly — random fluctuations will cross the line temporarily even when no real difference exists. Decide the sample size or duration in advance, run the test to completion (covering at least one, ideally two, full weekly cycles), and evaluate once at the end.

How many conversions do I need for this test to be trustworthy?

As a working rule, at least 100 conversions per variant before the normal approximation this calculator uses becomes dependable, and 200–300+ per variant to detect the modest 5–15% lifts most real tests produce. Small stores often lack the traffic to detect small effects — in that case test bigger, bolder changes where the true lift is large enough to find.

The calculator shows a big lift but low confidence — what does that mean?

It means your sample is too small to distinguish that lift from luck. A 20% lift on 30 conversions per side is statistically fragile; the same lift on 300 conversions per side is compelling. Low confidence is not evidence the variant failed — it is an instruction to keep collecting data or accept that the test is underpowered and move on.

Related calculators

Part of the Conversion & Email collection.