Quick definition: An error-spending function is a prespecified rule that allocates a study’s total type I error probability across interim analyses. It lets a sequential experiment inspect results at planned looks while preserving its overall false-positive rate under the stated design.
What is an error-spending function?
A conventional fixed-horizon test spends its entire false-positive budget, usually α = 0.05, at one final analysis. If a team repeatedly applies that same threshold while accumulating data, its chance of at least one false positive rises above 5%. An error-spending function replaces the repeated fixed-horizon test with a sequential design: it says how much of the total α may be used by each information time.
Let t be information fraction, from zero to one, rather than simply calendar time. An error-spending function α(t) gives cumulative type I error spent by the time an analysis has reached fraction t of planned information. At look j, the newly available amount is α(tj) − α(tj−1). A sequential-testing procedure converts those increments into critical boundaries, usually accounting for correlation between repeated test statistics.
Common functions and their assumptions
An O’Brien–Fleming-like function spends very little α early and most near the planned end. It demands strong early evidence, making it appropriate when premature launches are expensive and preserving near-fixed-horizon behavior matters. A Pocock-like function spends more evenly, yielding similar boundaries at each look and enabling earlier stopping more readily, but usually requiring a somewhat larger maximum sample size for comparable power.
Lan–DeMets alpha spending is commonly used because it can approximate these boundary families even when actual look times differ from the original schedule. This flexibility does not mean analysts can create arbitrary looks after seeing promising data. The total α, function family, primary endpoint, directionality, maximum information, and stopping action must be part of the analysis plan. Information fraction should reflect precision—often accrued sample size or inverse variance—not a dashboard’s elapsed days when traffic and outcome maturity vary.
The guarantees depend on the statistical model and operational protocol. The primary metric, assignment unit, analysis population, treatment definition, and analysis method must remain stable. If users are clustered, outcomes are delayed, or variance changes sharply, the information calculation and boundaries must reflect that design. The function controls a false-positive probability under its assumptions; it does not diagnose sample ratio mismatch, instrumentation errors, or interference between users.
Application to A/B testing
Error spending supports experiments in which waiting until a fixed end date is costly or unethical. A harmful reliability change can be stopped early; a clearly superior onboarding flow may be rolled out sooner; an inconclusive experiment can continue to maximum information. The method is particularly valuable when a product organization otherwise checks a p-value every day and reacts to whichever day looks favorable.
Before launch, define the primary estimand and practical decision criteria. For example: conversion among all eligible assigned users within seven days, a two-sided α of 0.05, 90% power for a 0.4-point effect, four information looks, and a O’Brien–Fleming-like spending function. Separately define efficacy, futility, and guardrail actions. Alpha spending addresses false positive control for the primary efficacy test; it does not automatically control multiple metrics, variants, segments, or changing analyses. The planning issues are related to multiple comparisons in A/B tests.
Use mature outcomes. If a seven-day conversion metric is analyzed before treatment users have had seven days, the apparent information can be misleading. Either delay each look until the cohort matures or use a validated time-to-event method. Record the data cut, information fraction, boundary, observed estimate, and decision for every look. A reproducible audit trail prevents a “planned” sequential design from becoming undocumented optional stopping.
Worked example: four planned looks
A team tests a checkout flow with a maximum of 100,000 eligible users and a two-sided α = 0.05. It plans looks at 25%, 50%, 75%, and 100% information using an O’Brien–Fleming-like spending function. Illustrative two-sided z-boundaries might be approximately 4.05, 2.86, 2.34, and 2.02; the precise values depend on the selected implementation and covariance structure.
At 25% information the estimated conversion lift is +0.8 points but z = 2.6. It is impressive on a dashboard but below the early boundary, so the efficacy rule says continue. At 75%, z = 2.46 and the estimate is +0.45 points; it crosses the precomputed boundary. The team rejects the null for the primary metric, then checks predeclared payment-error and page-latency guardrails before deciding whether a limited or full rollout is appropriate.
If the test had reached the final look with z = 1.8, it would not cross the final efficacy boundary. The correct statement is that the planned sequential test did not reject the null; it is not proof of no commercially important effect. Report the effect estimate and interval or compatible uncertainty measure, and compare remaining uncertainty with the minimum useful effect.
Interpretation and decision workflow
- Set maximum information, alpha, sidedness, primary metric, and the spending function before exposure.
- At each scheduled or justified information look, validate eligibility, allocation, data maturity, and metric computation.
- Calculate the sequential statistic and compare it with the boundary for that exact information fraction.
- Take the prespecified action: stop for efficacy, stop for harm or futility, or continue.
- Report all looks and boundaries, not only the one that crossed.
Crossing an efficacy boundary is evidence against the specified null under the sequential design; it is not a complete launch recommendation. The observed effect at a stopping boundary can be larger than the eventual true effect, especially for early stopping. Use shrinkage or follow-up measurement where the effect size drives a high-stakes commitment, and assess practical importance using absolute business units. Guardrail review follows the same principles described in primary and guardrail metrics.
Limitations and common mistakes
- Using calendar time as information. Delayed outcomes, unequal allocation, and variance changes mean elapsed time may not represent precision.
- Mixing boundary families. Critical values must come from the actual α function, looks, sidedness, and analysis model.
- Adding unplanned looks casually. Flexible timing needs a validated Lan–DeMets-style implementation and documentation.
- Changing the endpoint mid-test. A new metric or denominator creates a different test and can invalidate the nominal control.
- Ignoring effect-size exaggeration. Early stopped estimates need careful communication and may warrant confirmation.
- Calling it a cure for data quality. Sequential inference cannot restore causal validity to corrupted assignment or measurement.
Frequently asked questions about error-spending functions
Is error spending the same as alpha spending?
Alpha spending is the usual type I error case. “Error spending” can be broader, including beta spending for power or futility, but the term commonly refers to allocating alpha.
Can I look at the dashboard whenever I want?
You can monitor operational health continuously, but formal efficacy decisions should follow the validated sequential rule. Separate safety diagnostics from outcome-based winner declarations.
Does a sequential test need more users?
Often the maximum sample size is modestly higher than a fixed-horizon design to retain the same power, although the expected sample size can be lower when effects are large or harmful.
What is the difference between efficacy and futility boundaries?
An efficacy boundary supports rejecting the null. A futility boundary supports stopping because the chance of reaching a useful result is too low under a stated rule; it has different error considerations.
Can this be used with Bayesian monitoring?
Bayesian decision rules use different probability guarantees, but they still need a prespecified, simulated operational policy. Do not combine frequentist boundaries and posterior thresholds without understanding the resulting error behavior.
Summary
An error-spending function distributes a finite false-positive budget across planned sequential analyses. It enables disciplined early decisions only when information timing, endpoints, boundaries, stopping actions, and data-quality checks are specified and executed consistently.
Sources
- Lan and DeMets, Discrete Sequential Boundaries for Clinical Trials
- O’Brien and Fleming, A Multiple Testing Procedure for Clinical Trials
- NIST/SEMATECH e-Handbook: Sequential Sampling