Design·Glossary term

Fixed-Horizon Test

Fixed-Horizon Test A/B testing Reference guide

Fixed-Horizon Test is a concept used in experiment design & methodology.

Quick definition: A fixed-horizon test is an experiment that commits before launch to a sample size or information target, an outcome-maturity rule, and one primary final analysis, rather than repeatedly deciding from ordinary interim p-values.

What is a fixed-horizon test?

A fixed-horizon test is the standard design for a controlled experiment with a predefined endpoint. Before assignment begins, the team specifies how much valid information it needs, usually expressed as a number of eligible randomized units with mature outcomes. It then collects data to that target and performs the planned primary analysis once. The design’s familiar confidence intervals and p-values are calibrated for this one planned analysis.

“Fixed” does not mean inflexible in every operational detail. Traffic can arrive at an uneven rate, the calendar end date can shift because outcome windows have not matured, and an emergency harm response can stop exposure. It means that the confirmatory statistical decision does not move in response to favorable or unfavorable outcome trends. The horizon is set by the design, not by a dashboard.

This distinction matters because repeated peeking changes error rates. If a team tests the same null hypothesis every day and stops the first time an ordinary p-value falls below 0.05, the chance of eventually reporting a false positive exceeds 5%. A fixed-horizon analysis avoids that particular optional-stopping problem by waiting for the prespecified information target. Teams that need scheduled interim decisions should use a sequential testing method instead of relabeling repeated looks as fixed horizon.

What a valid fixed-horizon design specifies

A protocol starts with a decision, not a desired p-value. State the target population, randomization unit, control and treatment experiences, primary metric, observation window, and estimand. For example, the question may be whether a new activation checklist increases activation within 14 days per eligible randomized user by at least 0.5 percentage points without increasing support contacts above a defined guardrail threshold.

Next, choose a significance level, alternative hypothesis, target power, baseline estimate, minimum effect of interest, allocation ratio, and variance or conversion assumptions. These inputs yield a required sample or information size. The minimum detectable effect is not a claim that smaller effects are zero; it is the effect the test is designed to detect reliably.

Protocol componentWhy it mattersFailure if omitted
Primary metric and windowDefines the outcome and when it is completeMetric or maturity switching after results appear
Horizon or information targetSets the final evidence pointStopping when a chart looks favorable
Effect threshold and powerConnects sample size to a product decisionAn inconclusive test framed as no effect
Analysis populationPreserves the randomized comparisonPost-treatment filtering and selection bias
Guardrails and escalation rulesProtects users and operationsWaiting for a final analysis during clear harm

The planned horizon is normally based on mature outcomes, not raw assignments. A 28-day retention experiment has not reached its final information size merely because 20,000 users were assigned yesterday. Define a data cutoff that gives every included user the full window, specify late-event handling, and keep a separate count of assigned, exposed, and outcome-mature units.

Assumptions and validity conditions

Fixed-horizon inference assumes the protocol’s primary analysis was the analysis used to make the claim. Randomization must be correctly implemented, each arm must receive the intended experience, and outcome measurement must be comparable. The unit of analysis must reflect the unit of assignment and dependence. If stores are randomized, treating every customer as an independent treatment assignment understates uncertainty; if users return repeatedly, aggregate or model repeated observations according to the estimand.

The test also needs stable eligibility and definitions. A mid-test change to the activation event, traffic source, assignment logic, price, or identity rule can change the population or outcome. Document unavoidable changes, assess whether both arms were affected equally, and determine whether the original estimand remains meaningful.

Missing outcomes require a predeclared approach. If treatment causes more users to uninstall, become untrackable, or exit a workflow, excluding those users may hide a real treatment consequence. Analyze by assignment whenever feasible and report differential missingness. Check traffic splits and exposure logging; a sample ratio mismatch is a data-quality signal, not a nuisance to ignore.

Important: A fixed horizon protects the stated final comparison. It does not validate unplanned metric selection, segment hunting, broken randomization, or switching to a favorable cutoff.

When a fixed horizon is the right experiment design

Choose a fixed horizon when the team has a clear primary decision, a stable enough traffic forecast, and no need for routine early efficacy decisions. It is especially attractive for long-lag outcomes, modest expected effects, regulated or high-accountability decisions, and teams that need a simple auditable process. The interpretation is straightforward: at the prespecified endpoint, estimate the effect and compare its uncertainty with business thresholds and guardrails.

It is not always the most efficient design. For a high-risk change, waiting until the end is inappropriate if credible harm appears; an emergency safety process should allow pause or rollback. For a change where early success or futility would materially affect resources, a group sequential design can be appropriate, provided its boundaries and analysis are preplanned. The choice is between valid decision frameworks, not between rigor and speed.

Worked scenario: a new account-verification flow

A financial app tests a shorter account-verification flow. The primary metric is successful verification within seven days among eligible randomized applicants. Fraud-confirmed accounts, manual-review rate, and support contacts are guardrails. Based on a baseline verification rate of 62%, a minimum useful lift of 1.5 percentage points, 90% power, and two-sided 5% significance, the team plans 28,000 mature applicants per arm. It assigns applicants 1:1 and fixes the final analysis after every included applicant has had seven days for the outcome to mature.

Daily operational monitoring shows traffic, assignment counts, exposure, payment-provider errors, and severe fraud alarms. It does not use the accumulating verification p-value to choose a finish date. On day 20, a new fraud rule is deployed globally. The team records the change, checks that both arms were affected, and includes calendar controls only if they were defined as a sensitivity analysis. If the fraud system had been enabled only for treatment, the original causal comparison would need reassessment rather than an automatic final calculation.

At the horizon, verification improves by 1.7 points with an interval of 0.6 to 2.8 points. Manual review declines, but fraud-confirmed accounts have an interval that includes a potentially unacceptable increase. The correct decision is not an immediate unrestricted rollout based on the primary p-value. The team holds the flow to a monitored staged rollout while gathering the safety outcome, because the protocol treats the fraud guardrail as decision-critical.

Analysis and decision process

  1. Write the product decision, causal estimand, population, assignment unit, primary metric, observation window, and guardrails.
  2. Set the minimum useful effect, alpha, power, allocation, baseline assumptions, and mature-information target; assess the plan with realistic traffic and variance scenarios.
  3. Predefine eligibility, assignment persistence, exposure criteria, missing-data handling, metric calculations, transformations, and primary model.
  4. Validate assignment, tracking, identity rules, and dashboard counts in an A/A or dry run before interpreting treatment outcomes.
  5. During collection, monitor operational health and predefined safety signals; avoid final-hypothesis decisions from repeated ordinary looks.
  6. Lock the mature dataset at the horizon and run the planned analysis, reporting absolute values, effect estimate, interval, p-value if relevant, and every guardrail.
  7. Compare the result with practical thresholds, implementation cost, risk, and a rollback plan; publish deviations and the rationale for ship, iterate, or stop.

Sample calculations should be transparent and revisited only through a documented protocol amendment. A lower-than-expected baseline rate or greater outcome variance can mean the original horizon lacks power, but continuing until significance is not a valid remedy. Recalculate the information requirement under a proper amendment or treat an underpowered result as inconclusive. For the mechanics, see how to calculate sample size.

Limitations and common errors

  • Peeking and stopping: ending early after an unadjusted favorable result inflates false-positive risk.
  • Calendar thinking: stopping before outcomes mature can answer a different question.
  • Underpowered null claim: a non-significant result does not establish no meaningful difference without adequate precision.
  • Metric substitution: promoting a secondary metric because the primary was inconclusive is exploratory unless planned.
  • Ignoring practical significance: a precisely estimated tiny effect may not repay implementation cost.
  • Unmanaged risk: a fixed horizon is not permission to continue a harmful experience.

Frequently asked questions

Can we look at the data before the fixed horizon?

You can monitor assignment, logging, exposure, and safety. Avoid using repeated ordinary significance tests to make efficacy decisions; use a prespecified sequential method if interim efficacy decisions are required.

What if traffic is slower than planned?

Extend calendar time until the mature-information target is reached, revise the plan transparently, or accept less precision. Do not lower the target solely because the result is inconvenient.

Can a fixed-horizon test stop for harm?

Yes. Ethical and operational safety overrides are appropriate, but document the rule and recognize that the final efficacy analysis may need a different interpretation.

Does a fixed horizon guarantee a valid result?

No. It addresses one form of optional stopping. Randomization, measurement, population definition, exposure, and analysis choices still determine whether the estimate is credible.

Is a p-value the rollout decision?

No. It is one evidence summary under a model. Rollout also depends on effect size, interval, guardrails, user risk, implementation quality, and the minimum value needed to justify change.

Summary

A fixed-horizon test commits to a mature-information target and one planned primary analysis, preventing ordinary dashboard peeking from determining the endpoint. It is a clear, robust default when decisions do not need routine interim efficacy actions. A credible result still requires clean randomization and measurement, a practical effect threshold, outcome maturity, and a decision process that weighs guardrails alongside statistical evidence.

Sources