Worked examples
Promising but not proven
| Visitors — variant A | 5000 |
| Conversions — variant A | 150 |
| Visitors — variant B | 5000 |
| Conversions — variant B | 180 |
| Lift (B vs A) | 20.0% |
| Z-score | 1.68 |
| Confidence | 90.7% |
A 20% lift at ~91% confidence — keep the test running to reach 95%.
Clear winner
| Visitors — variant A | 8000 |
| Conversions — variant A | 240 |
| Visitors — variant B | 8000 |
| Conversions — variant B | 320 |
| Lift (B vs A) | 33.3% |
| Z-score | 3.44 |
| Confidence | 99.9% |
A 33% lift with z ≈ 3.4 — beyond 99.9% confidence; ship variant B.
Frequently asked questions
What confidence level should I require before calling a winner?
95% is the standard bar — it means a difference this large would appear by chance in fewer than 1 in 20 no-difference tests. For low-risk changes like button copy, some teams accept 90%; for pricing or checkout changes where a wrong call is expensive, hold out for 95% or higher plus a sensible minimum run time of two full business cycles.
Why should I not peek at results and stop early?
Checking repeatedly and stopping the moment confidence crosses 95% inflates your false-positive rate badly — random fluctuations will cross the line temporarily even when no real difference exists. Decide the sample size or duration in advance, run the test to completion (covering at least one, ideally two, full weekly cycles), and evaluate once at the end.
How many conversions do I need for this test to be trustworthy?
As a working rule, at least 100 conversions per variant before the normal approximation this calculator uses becomes dependable, and 200–300+ per variant to detect the modest 5–15% lifts most real tests produce. Small stores often lack the traffic to detect small effects — in that case test bigger, bolder changes where the true lift is large enough to find.
The calculator shows a big lift but low confidence — what does that mean?
It means your sample is too small to distinguish that lift from luck. A 20% lift on 30 conversions per side is statistically fragile; the same lift on 300 conversions per side is compelling. Low confidence is not evidence the variant failed — it is an instruction to keep collecting data or accept that the test is underpowered and move on.
Related calculators
- Conversion Rate Calculator
- Cart Abandonment Cost Calculator
- Email ROI Calculator
- Break-even ROAS Calculator
- All calculators
Part of the Conversion & Email collection.