Numeric metrics

T-Test Calculator

Compare means with Welch’s independent-samples t-test or a paired t-test from summary statistics.

Why use this calculator

Separate a shift in averages from random variation

A t-test compares a difference in means with the variation and sample sizes behind it. It is a common analysis for numeric experiment metrics such as revenue per assigned user, time measured for every randomized unit, or a continuous quality score.

Compare means

Is the mean difference larger than noise?

Use summary statistics from independent groups, or statistics for the within-pair differences.

Test type

Control

Treatment

P-value

—

t statistic · degrees of freedom

—

Mean difference

—

Standard error

—

Degrees of freedom

—

What does a t-test measure?

A t-test standardizes the observed difference in means by its standard error. The resulting p-value describes how unusual a difference at least this large would be if the relevant population means were equal.

It does not measure business impact. Report the mean difference in the metric’s original units and its confidence interval alongside the p-value, then compare the range with the smallest effect that would change your decision.

Choosing an independent or paired t-test

Independent groups: use Welch’s t-test by default

Use this mode for separately randomized control and treatment groups. Welch’s version does not assume the two groups have identical variance, so it is generally a safer default than the pooled-variance t-test.

Paired data: analyze the within-pair difference

Use paired mode when every treatment observation is naturally matched to one control observation, such as before-and-after measurements on the same person. Do not use it for ordinary independently randomized A/B-test users.

Define the metric for all randomized units

For experimentation, include every assigned unit under the pre-specified metric definition. Restricting analysis to people who completed a treatment-affected step can introduce selection bias.