Design·Glossary term

Peeking

Peeking A/B testing Reference guide

Peeking is a concept used in experiment design & methodology.

Quick definition: Peeking is examining an experiment’s accumulating results before the prespecified analysis and allowing what is seen to influence a decision.

What is Peeking?

Peeking is examining an experiment’s accumulating results before the prespecified analysis and allowing what is seen to influence a decision. A fixed-horizon A/B test has one confirmatory analysis at its planned information size. Repeatedly applying an ordinary 0.05 threshold after each dashboard refresh creates many chances for random variation to look persuasive. A valid alternative specifies interim looks and decision boundaries in advance, or uses an always-valid method whose guarantee accommodates continuous review.

Its practical value depends on using it as part of a complete measurement system: a stable population, documented eligibility, reliable assignment and exposure records, an outcome that represents user value, and an analysis that is specified before the result drives a choice. The label alone does not establish causality or make a decision safe.

Method and design

A fixed-horizon A/B test has one confirmatory analysis at its planned information size. Repeatedly applying an ordinary 0.05 threshold after each dashboard refresh creates many chances for random variation to look persuasive. A valid alternative specifies interim looks and decision boundaries in advance, or uses an always-valid method whose guarantee accommodates continuous review.

Design choices should be made before meaningful outcome differences are available. Teams should document data sources, cutoff times, transformations, allocation rules, and the exact population included in the estimate. These details make the comparison reproducible and reveal whether an apparent improvement could instead reflect a change in composition, measurement, or timing.

Assumptions and checks

The monitoring plan must define the primary endpoint, information timing, population, maximum sample, outcome maturity, and rule for efficacy, futility, and safety. Assignment, exposure logging, and metric definitions must remain comparable at every look.

Assumptions are not boilerplate. Test what can be tested with balance tables, pre-period plots, logging audits, sensitivity analyses, placebo checks, and outcome-quality reviews. For assumptions that cannot be directly tested, explain why they are plausible and how a violation would change the decision. The appropriate response to a failed check is to investigate, revise the claim, or collect better evidence—not to suppress the diagnostic.

Scenario and decision workflow

A checkout team sees a favorable conversion result after three days and wants to ship. The primary metric has a seven-day maturation window, so only part of the assigned population can contribute a complete outcome. The team waits for the scheduled sequential review, evaluates mature users and guardrails together, and records whether the prespecified efficacy or harm boundary was crossed.

The team then compares the observed effect with its minimum useful threshold and makes the action proportional to uncertainty. A robust positive result may support a controlled rollout; an ambiguous result may support more collection or a redesign; evidence of harm can justify stopping or limiting exposure. In every case, the report distinguishes the observed estimate from assumptions about future scale.

Using Peeking in A/B testing

In product experimentation, the method should serve a decision rather than decorate a report. Begin with the eligible population, unit of randomization, exposure event, primary metric, minimum useful effect, and guardrails. State whether the goal is a broad rollout, a targeted policy, a learning result, or a safety decision. That scope determines which evidence is relevant and prevents a technically correct analysis from answering a different question.

Keep the primary analysis separate from diagnostics and exploration. A concurrent control group remains the most direct protection against time trends and changing traffic mix when randomization is feasible. Before interpreting a result, verify assignment balance, sample-ratio behavior, exposure logging, metric maturity, and exclusions. A sophisticated design cannot compensate for a missing treatment event or a denominator that changed differently by arm.

Practical workflow

  1. Write the causal question and the operational decision in plain language.
  2. Define eligibility, assignment unit, treatment, exposure, primary outcome, guardrails, and analysis window before launch.
  3. Document the method-specific assumptions and how they will be checked.
  4. QA randomization, data lineage, outcome maturity, and population balance while the experiment runs.
  5. Run the prespecified analysis, then report diagnostics and clearly label exploratory work.
  6. Compare the estimate and interval with a practical decision threshold, not only a significance label.
  7. Record deviations, rollout conditions, and follow-up monitoring so the conclusion can be audited.

How to interpret results

An estimate is a range of plausible effects under the design’s assumptions, not a guarantee of future business impact. Report the absolute outcome in each condition, the effect on a stated scale, a measure of uncertainty, the eligible denominator, and the period represented. Compare the interval with a minimum useful effect and with relevant guardrails. A statistically detectable but tiny change can be a poor rollout choice; an imprecise result may justify more data rather than a binary verdict.

Interpretation also requires transportability. The result applies first to the population, implementation, traffic conditions, and measurement rules used in the study. Changes in audience targeting, pricing, seasonality, product dependencies, or operational capacity can make a later rollout behave differently. Preserve a contemporaneous control where possible and monitor the release rather than assuming an experimental average will remain unchanged.

Limitations and common mistakes

Treating every dashboard view as harmless; stopping a winner while extending a loser; calling immature outcomes final; adding a metric after a favorable trend; and describing a data-driven stop as if it were planned.

  • Changing the question after results arrive: define the decision and primary outcome before seeing a favorable pattern.
  • Ignoring the randomization unit: account-, store-, geo-, and time-based assignment need cluster- or design-aware uncertainty.
  • Using immature data: delayed revenue, retention, and refunds need completed observation windows.
  • Overlooking multiplicity: many metrics, segments, looks, and transformations create extra false-positive opportunities.
  • Hiding exclusions or deviations: document them because they can change the target population and conclusion.
  • Confusing precision with validity: a narrow interval around a biased estimate is still misleading.

Frequently asked questions

Is this method a replacement for randomization?

Usually no. When user- or account-level randomization is feasible and does not create unacceptable interference, it is generally the clearest way to identify a causal effect. This method either improves a randomized design or provides a structured alternative when randomization is constrained.

What should be prespecified?

Prespecify the decision, population, outcome, analysis window, primary comparison, key assumptions, stopping or review rule where relevant, and handling of missing data and multiple comparisons. Exploratory analysis remains useful when it is labeled and validated appropriately.

How much data is enough?

Enough data depends on the baseline variability, desired minimum detectable effect, allocation, correlation or clustering, and the precision needed for the decision. Use a planning calculation and assess achieved uncertainty rather than relying on a universal sample-size target.

Can a non-significant result prove no effect?

No. It means the data and method did not establish the stated contrast at the selected threshold. Inspect the interval: it may rule out a meaningful gain, remain compatible with both gain and loss, or support a practical equivalence conclusion only if that was designed and analyzed explicitly.

What belongs in the experiment report?

Include the decision context, protocol, population, assignment and exposure details, metric definitions, data cutoff, checks, effect estimates, uncertainty, guardrails, deviations, limitations, and the final action. A reader should be able to tell what was planned, what occurred, and what the evidence can support.

Summary

Peeking is examining an experiment’s accumulating results before the prespecified analysis and allowing what is seen to influence a decision. A fixed-horizon A/B test has one confirmatory analysis at its planned information size. Repeatedly applying an ordinary 0.05 threshold after each dashboard refresh creates many chances for random variation to look persuasive. A valid alternative specifies interim looks and decision boundaries in advance, or uses an always-valid method whose guarantee accommodates continuous review. Reliable use requires a clearly stated decision, transparent assumptions, high-quality data, and an analysis that matches the way the experiment was actually run.

Sources