Experiment planning

Sample Size Calculator

Plan a two-group experiment around the smallest effect worth detecting — for conversion rates or numeric metrics.

Why use this calculator

Set the test size before the result can influence the decision

A sample size calculator turns a business-relevant minimum detectable effect into a concrete experiment plan. It helps avoid tests that finish with too little evidence to detect a real win, while keeping the required traffic proportional to the decision you need to make.

Plan a two-group test

How many observations do you need?

Use the same metric definition and unit of randomization that you will use in the final analysis.

Metric type

Test settings

Required sample

Control

—

assigned users

Treatment

—

assigned users

Total

—

assigned users

Two-sided, fixed-horizon normal approximation. Plan the same analysis you intend to run.

How to use this estimate

For conversion metrics, enter the rate among all assigned users — including zeroes for people who do not convert. For numeric metrics, use a standard deviation estimated at the same unit of analysis as the final test.

The smaller allocation is not a reason to wait for a balanced sample. The estimate above calculates the observations required in each arm at the allocation you selected.

How to calculate sample size for an A/B test

An A/B test sample size calculation answers a practical planning question: how many assigned users are needed to reliably detect the smallest effect that would matter? The answer depends on the metric, its baseline variability, the minimum detectable effect (MDE), your significance level, and the power you want.

Start with the decision, not the traffic

For a conversion rate, choose a baseline and a relative lift worth acting on. For example, a 5% conversion rate and a 10% relative MDE means planning to detect a change from 5.00% to 5.50%. For a numeric metric, use the baseline mean and standard deviation from comparable historical data, then enter the smallest absolute difference that matters.

Power and alpha set the evidence threshold

Alpha controls the planned false-positive rate; 5% is a common two-sided default. Statistical power is the chance of detecting the chosen effect if it is real; 80% is a common starting point. A smaller MDE, lower alpha, or higher power requires more observations.

Unequal traffic changes each arm’s requirement

A 50/50 split is efficient, but experiments sometimes allocate less traffic to a risky treatment. In that case, the calculator increases the larger arm so the smaller arm still has enough observations. If you enter daily eligible traffic, the result also estimates how long the planned allocation will take to fill.

Use a specialized method when needed. This calculator is for independent two-group, fixed-horizon tests. Clustered or paired designs, sequential methods, and ratio metrics such as average order value can require a different variance model and should be planned with the same method used for analysis.