How to design a rigorous A/B test that produces trustworthy results.
1. Define the hypothesis clearly
2. Choose the primary metric
3. Calculate sample size
required_sample_size_proportions() for conversion metricsrequired_sample_size_means() for continuous metrics4. Define guardrail metrics
5. Plan for multiple testing
| Pitfall | Why It's Bad | Fix |
|---|---|---|
| Peeking at results | Inflates false positive rate | Use sequential testing |
| Under-powered test | High chance of missing real effects | Calculate sample size first |
| Too many variants | Dilutes traffic, extends runtime | Max 3-4 variants |
| Wrong randomization unit | Users see inconsistent experience | Randomize by user, not session |
| Novelty effect | Initial lift fades over time | Run for 2+ weeks minimum |
| Day-of-week effects | Weekend vs weekday behavior differs | Run for complete weeks (7, 14, 21 days) |
| Baseline Rate | MDE (absolute) | Approx. N per group |
|---|---|---|
| 5% | 1pp | ~7,000 |
| 5% | 2pp | ~2,000 |
| 10% | 1pp | ~14,000 |
| 10% | 2pp | ~4,000 |
| 20% | 2pp | ~6,000 |
| 20% | 5pp | ~1,000 |
| 50% | 5pp | ~1,600 |
These assume alpha=0.05, power=0.80, two-sided test.
Get the full A/B Testing Statistical Framework and unlock everything.
Get the complete guide with every chapter unlocked, including code samples, diagrams, and best practices.
Access all interactive tools with complete data, all workload profiles, and the full scenario library.
Downloadable source code, configuration files, and working examples from every chapter.
Free updates for life. Every new chapter, tool, and improvement included.