Quick definition: A holdout group is a deliberately withheld, randomly assigned group that does not receive a product, marketing, or policy change, so its outcomes can estimate the change’s incremental impact.
What is a holdout group?
A holdout group is a comparison population kept on a reference experience while another population receives an intervention. It provides a live answer to the counterfactual question: what would these eligible users, customers, accounts, stores, or markets have done if the change had not been available? When assignment is random and the groups operate concurrently, differences in outcomes can be attributed more credibly to the intervention than to unrelated changes in demand or environment.
The term is sometimes used loosely for any users excluded from a campaign. In rigorous experimentation, exclusion alone is not enough. A useful holdout is defined before outcomes, has a documented eligibility rule and assignment method, receives a specified reference experience, and is measured with the same outcome definitions as treatment. A random holdout differs from customers who happen not to adopt a feature; non-adopters may have different needs, awareness, or opportunity.
A holdout group overlaps with a control group, but the emphasis is often different. “Control” commonly describes one arm of a bounded A/B test. “Holdout” frequently describes a group withheld during a broad rollout, ongoing campaign, personalization policy, or long-term measurement program. The design principles are the same: a stable comparison, comparable measurement, and careful management of spillover.
How holdouts fit experimental design
First define the decision. A holdout can answer whether a new recommendation system creates incremental revenue, whether an email program changes retention beyond natural behavior, or whether a fully rolled-out feature still has durable value. It should not exist merely because an experimentation platform has a holdout setting. Withholding access creates a cost, so the expected learning must justify the lost benefit, operational complexity, and any fairness implications.
Set eligibility before assignment. For an upgrade campaign, eligibility might mean paid-plan prospects in supported countries who have not opted out of email. Assign at customer or account level if people can receive messages on multiple devices or through several channels. Persistent assignment prevents a customer from moving between treatment and holdout, which would blur the contrast. Use a randomization key that is available before treatment and audit the planned allocation with sample-ratio checks.
Define treatment and holdout experiences precisely. The treatment may receive a new recommendation model; the holdout may receive the prior model, a no-recommendation experience, or an established default. “No intervention” can be misleading if other teams, support agents, or automated systems still contact the holdout. Specify channel rules, suppression logic, version changes, and the period of protection. Log assignment and actual exposure separately, because intended assignment is not proof that a message was delivered or a feature rendered.
Incremental outcome: estimated incremental value = outcome rate in treatment − outcome rate in holdout, multiplied by the number of eligible units expected under rollout. Estimate uncertainty before translating the difference into revenue or cost.
Choose outcomes that represent value, not merely activity caused by the intervention. For a retention email, open rate is a delivery diagnostic, not incremental retention. A reasonable primary metric might be retained paid accounts after 60 days, with unsubscribe rate, complaints, discount cost, and support contacts as guardrails. Let outcome windows mature equally for both groups. Reading a 30-day result after ten days for one arm is not a valid comparison.
Practical scenario: a persistent recommendation holdout
An ecommerce company rolls out a new recommendation model on product pages. Offline modeling suggests that it should increase order value, but the team wants to know its ongoing causal contribution after launch. It randomly assigns 5% of eligible logged-in customers to a persistent holdout that sees the prior ranking; 95% receive the new model. The small allocation limits withheld opportunity while preserving enough information for a planned, longer measurement period.
The primary outcome is contribution margin per eligible customer over 28 days. Secondary measures include conversion, units per order, return rate, page latency, and diversity of products purchased. The design excludes employees and known bots before assignment, records inventory availability, and preserves the same ranking page architecture in both arms. Because customers can share links and inventory is limited, the team also watches for interference: treatment demand could alter availability and therefore the holdout’s outcomes.
After eight weeks, treatment has a positive margin estimate, but the interval is wide and return rate rises for a category with high fulfillment cost. The team does not simply remove the holdout because the dashboard looks favorable. It checks allocation, exposure, catalog changes, and whether the planned information target has been reached. It may continue collecting data, refine the model for the problematic category, or accept a smaller but more credible margin estimate. A persistent holdout supports learning after rollout; it does not eliminate the need for an ethical decision about who is excluded.
Holdout decision workflow
- State the causal question. Identify the incremental effect and the action that depends on it.
- Assess whether withholding is acceptable. Do not withhold essential service, safety protection, or a known necessary benefit without appropriate review.
- Set population and assignment. Choose pre-treatment eligibility, a stable unit, allocation, and duration that meet the decision’s precision needs.
- Specify both experiences. Document exactly what treatment receives and what the holdout continues to receive, including channels and fallbacks.
- Predefine outcomes and maturity. Select primary value, guardrails, attribution window, exclusions, and statistical monitoring rules.
- Validate implementation. Check assignment persistence, exposure records, suppression rules, and sample-ratio mismatch before interpreting lift.
- Make a proportionate decision. Combine the estimate, uncertainty, costs, risks, and reversibility; record why the holdout is retained, changed, or ended.
Allocation is a trade-off, not a default. A 50/50 split maximizes precision for a two-arm comparison when exposure costs are comparable. A 95/5 rollout holdout provides less precision but may be appropriate when treatment is already believed beneficial, the treatment is expensive to withhold, or the question concerns large long-run effects. Calculate sample size for the actual ratio. An arbitrarily tiny holdout can create a false sense of measurement while making material effects impossible to distinguish from noise.
Limitations and common mistakes
A holdout estimates the effect of the tested delivery under the tested conditions. It cannot prove that the intervention will work for customers excluded from eligibility, in a new season, or after surrounding systems change. Long-running holdouts are particularly exposed to product releases, changing inventory, identity merges, and altered customer behavior. Version the experiences and annotate major changes so the final comparison remains understandable.
- Non-random exclusions: comparing adopters with non-adopters introduces selection bias rather than a counterfactual.
- Holdout leakage: users receive treatment through another channel, a shared account, cached content, or support intervention.
- Moving users between arms: reassignment destroys a clean cumulative contrast unless the design explicitly models it.
- Measuring vanity metrics: impressions and clicks may rise even when net value does not.
- Ignoring interference: marketplace inventory, referrals, teams, and social effects can make arms affect each other.
- Keeping a holdout indefinitely by habit: review its cost, ethical basis, and continuing information value.
Holdouts also require governance. Teams should know who owns the allocation, who can override it in an incident, how consent and communication rules apply, and what stop conditions protect users. If an intervention becomes clearly necessary for safety or contractual service, the holdout should not be preserved to improve statistical power. The design serves responsible decisions, not the other way around.
Frequently asked questions
Is a holdout group always a control group?
It serves the control function when it is a valid concurrent comparison. “Holdout” often highlights that the group is withheld during a rollout or ongoing program.
How large should a holdout be?
Use a planned sample-size calculation based on baseline variability, minimum useful effect, allocation, and outcome window. Smaller is not automatically safer if it makes the result unusable.
Can a holdout receive the old experience?
Yes. The reference may be the previous model, a default policy, or no feature, as long as it answers the actual decision question and is documented.
When should a persistent holdout end?
End or revise it when its learning value no longer justifies withholding, its reference experience changes materially, or ethics, safety, and business constraints require full access.
What if treatment leaks into the holdout?
Measure and investigate the leakage. It often dilutes the estimated difference; redesign delivery or estimate a clearly defined assignment effect rather than claiming a pure exposure effect.
Summary
A holdout group preserves a randomized live comparison during a test, rollout, or ongoing program. Define it before assignment, specify both experiences, measure meaningful outcomes with mature windows, and balance the value of learning against the cost and ethics of withholding treatment.
Sources
- Control group glossary
- Experiment design glossary
- AB Labz: Primary versus guardrail metrics
- AB Labz: Sample-ratio mismatch in A/B testing