Effect Size Calculator
Translate an observed or planned difference into a standardized effect size for numeric metrics or conversion rates.
Why use this calculator
Put a difference in context
A raw difference is difficult to compare when a metric’s scale or variability changes. This effect size calculator standardizes the comparison, so you can distinguish a statistically detectable movement from a difference that is likely to matter in practice.
Which effect size should you use?
Cohen’s d standardizes a difference between two numeric means by their pooled standard deviation. Hedges’ g, shown in the result details, applies a small-sample correction.
Cohen’s h is a standardized difference between two proportions. It is useful alongside the percentage-point difference and relative lift, especially when baseline conversion rates are small.
How to interpret effect size in A/B testing
An effect size answers a different question from statistical significance. A p-value asks whether an observed difference would be unusual under a no-effect model. An effect size describes how large that difference is, after accounting for the metric’s variability or scale.
Cohen’s d for numeric metrics
Use Cohen’s d when comparing means such as revenue per assigned user, time spent, or the number of support contacts. A d of 0.5 means the mean difference is half of one pooled standard deviation. Labels such as small, medium, and large are useful orientation, but they do not replace a product-specific decision threshold.
Cohen’s h for conversion rates
A one-percentage-point conversion change does not have the same practical meaning at every baseline. Cohen’s h standardizes the difference on the arcsine scale, making it more comparable across rates. Read it alongside the raw percentage-point change and relative lift shown by the calculator.
Use effect size when planning, too
Standardized effect size helps connect historical variation with the minimum detectable effect in a sample size calculation. Before running a test, decide what outcome would be meaningful for customers or the business; after the test, compare that threshold with the observed effect and its uncertainty.