A/B Test Sample Size and Statistical Power Calculator
Plan concept, ad and pricing tests: sample size per cell, minimum detectable effect or statistical power for percentages and averages, with unequal cells.
inquiry@globalvoxpopuli.com · globalvoxpopuli.com
How to use the A/B test calculator
- Choose whether the outcome is a percentage (e.g. purchase intent, ad recall) or an average (e.g. liking on a 0–10 scale).
- Choose what to solve for. Sample size tells you how many completes each cell needs. Detectable gap tells you the smallest difference a fixed budget can pick up. Power tells you how likely an existing design is to detect a given gap.
- Enter the baseline result and the smallest gap that would change your decision.
- Keep 5% significance and 80% power unless you have a reason to change them.
- Enter the number of cells if you test several concepts or ads against one control.
What power means
Power is the chance that your test will show a real difference as significant. At 80% power, if concept B really scores 5 points higher than concept A, your test will flag it 8 times out of 10. The other 2 times, you would wrongly conclude there is no difference.
The significance level controls the opposite mistake: declaring a difference when there is none. Most research teams accept 5% for false positives and 20% for false negatives (80% power).
Monadic concept tests, ad tests and tracking waves
The calculator applies to any comparison of two independent samples:
- Monadic concept or pack tests: each respondent sees one concept, and cells are compared.
- Ad pre-tests: exposed vs control cells.
- Price tests: one price per cell.
- Tracker waves: wave 1 vs wave 2, with fresh samples each wave.
Sequential monadic designs, where each person rates several concepts, need fewer respondents because each person is their own control. This calculator gives a conservative (safe) answer for them.
Worked example
An FMCG client tests four new concepts monadically against its current one. Current top-2-box purchase intent is 22%. A 6-point lift would justify launch.
- At 5% significance and 80% power: 817 completes per cell
- Five cells in total: 4,085 completes
If the budget allows only 300 per cell, the detectable gap is about 10 points, so smaller wins would go unnoticed. Press Example to load the case, then switch to Detectable gap.
Frequently asked questions
What is a minimum detectable effect?
It is the smallest true difference your test can detect with the chosen power and significance. Anything smaller may be real but will often show as "not significant".
Should I use equal cell sizes?
Equal cells are most efficient for a fixed total. Use a larger control cell when it is compared with several test cells, often √(number of test cells) times larger.
Why is my required sample so large?
Sample size grows with the square of precision. Detecting a 3-point gap needs about four times the sample of a 6-point gap. Decide which gap would actually change the decision.
Is this the same as website A/B testing?
The statistics are the same. Survey tests usually have smaller samples and larger effects than website conversion tests, so they reach decisions with hundreds rather than tens of thousands of respondents.
Need the respondents, not just the numbers?
Global Vox Populi runs quantitative and qualitative fieldwork in 170+ countries through its own proprietary panels of consumers, B2B and IT decision-makers, physicians, nurses, patients, caregivers and payers. ISO 9001, ISO 20252 and ISO/IEC 27001 certified and HIPAA compliant, following ESOMAR and Insights Association guidelines.
Check feasibilityLink to this tool
Free to use in proposals, courses and articles. If it helped, a link back is appreciated:
<a href="https://globalvoxpopuli.com/tools/ab-test-sample-size-calculator/">A/B Test & Power Calculator</a> by Global Vox Populi