Experiment planning

Power Calculator

Check whether your planned sample has enough sensitivity to detect an effect worth acting on.

Why use this calculator

Find out what your current plan can actually detect

Statistical power is the probability that an experiment will detect its planned effect when that effect is real. Use this calculator before launch to check whether a fixed amount of traffic is enough, or after planning a sample size to confirm the trade-off you chose.

Evaluate a two-group test

What is the planned power?

Enter the observations available per arm and the smallest effect you want the test to detect.

Metric type

Planned power

—

—

— total assigned users

Two-sided, fixed-horizon normal approximation for independent groups.

What does statistical power mean?

Power is not the probability that your result will be significant. It is the long-run probability that the test rejects the no-effect hypothesis when the specific effect you planned for is present. At 80% power, roughly one in five tests would still miss that effect because of sampling noise.

More observations, a larger minimum detectable effect, a higher baseline conversion rate, and a less stringent alpha can all increase power. Use the same metric definition and analysis method here that you intend to use for the final experiment result.

How to use a power calculator for an A/B test

A power calculator is useful whenever traffic, duration, or risk is fixed before an experiment starts. Instead of asking only “how many users do we need?”, it asks whether the users you can realistically expose give the test enough chance to find a meaningful effect.

Check feasibility before launch

If planned power is low, a non-significant result will be difficult to interpret: it may indicate no meaningful difference, or simply insufficient data. Consider a longer duration, more traffic, a larger MDE, or a more precise metric before committing to the test.

Do not calculate power from the result you already saw

Post-hoc power based on the observed effect mostly restates the p-value. For decision-making, define the effect worth detecting in advance and calculate power from that planned effect.