Quick definition: A geo-lift test measures the incremental outcome from advertising or another market-level intervention by comparing randomized treated geographies with similar untreated geographies over the same period.
What is a geo-lift test?
A geo-lift test is a geographic experiment designed to estimate lift: the change in an outcome caused by an intervention beyond what would have happened without it. Most often, it measures incremental sales, sign-ups, profit, or visits produced by media spend. The treatment is delivered to selected geographic markets, while control markets retain the baseline plan. The post-test difference, adjusted according to the prespecified design, is evidence about causality rather than a platform’s attribution credit.
“Lift” must be defined in an outcome and scale. It may mean 400 additional weekly orders per market, a 3% relative increase in net revenue, or $1.30 of incremental gross profit per advertising dollar. It does not mean clicks, impressions, or conversions credited by a last-click model. Those measures may be useful diagnostics, but a geo-lift test asks whether total business outcomes changed relative to a valid counterfactual.
Geo-lift testing is a specialized use of a geo experiment. The name emphasizes incrementality and media measurement, while the validity requirements remain: a defined market unit, random or defensible assignment, concurrent controls, faithful delivery, stable measurement, and cluster-aware analysis.
Methodology and design choices
Start with the causal question. “What incremental gross profit does increasing paid social by 30% generate over six weeks among eligible markets?” is better than “Does paid social work?” Specify the media channels, spend increment, bidding, audience, creative, geographic boundary, flight dates, and baseline activity. Also specify the outcome data source, return and cancellation treatment, geographic attribution rule, and any lag or washout period.
Select markets from a universe where the campaign can be delivered and the outcome is observed consistently. Use historical sales level, trend, seasonality, population, channel availability, and operational conditions to form matched pairs or strata. Randomly assign treatment within the pairs. Pairing reduces variance by comparing places that behaved similarly before the intervention, but it cannot justify changing partners after the outcome is known.
| Question | Design answer | Failure if omitted |
|---|---|---|
| What changes? | Incremental media or policy, with delivery specifications. | The treatment is too vague to reproduce. |
| Where? | Non-overlapping market boundaries and assignment map. | Media and customers leak across groups. |
| What outcome? | Predeclared sales, profit, registrations, or other business metric. | Attribution metrics replace causal outcomes. |
| Compared with what? | Concurrent matched controls or randomized allocation. | Seasonality and broad trends confound lift. |
Plan power at the level of markets or matched pairs. Customer rows are correlated within geography and are not independent randomized units. Historical simulation is often the most credible planning method: repeatedly assign pseudo-treatment markets in pre-period data, calculate the planned estimator, and assess uncertainty and detectable lift. Include plausible delivery variance and contamination. If the available number of geographies cannot distinguish a commercially meaningful gain from noise, the test is not rescued by a sophisticated dashboard.
Assumptions and validity checks
Randomized assignment makes unmeasured differences less concerning on average, but the implementation must preserve that assignment. Verify actual spend, impression geography, frequency, auction conditions, creative availability, and whether other channels or local teams changed activity. A control market receiving retargeting from a treatment audience is not untouched. A treated market with stockouts tests a combined media-and-inventory condition.
Spillover is a central risk. People cross market borders, consume national media, search from travel locations, share offers, and purchase online. Use sufficiently separated markets, buffer areas, exclusion zones, delivery-address outcomes, or an analysis that explicitly measures exposure spillover. Do not assume that a platform’s geo-targeting setting proves clean isolation. The relevant question is whether the treatment-control contrast differs enough to identify the campaign effect.
Pre-period outcomes should be examined for stability and balance. Large differences do not necessarily invalidate randomization, but they reduce precision and can reveal a bad match. The test also needs a rule for local shocks—storms, store closures, competitor openings, pricing changes, tracking outages, or a concurrent experiment. Define exclusions and sensitivity analyses before reading the treatment effect.
A/B testing application
Geo-lift follows A/B-testing principles at a different unit of randomization. “A” is the baseline media plan in control markets; “B” is the incremental campaign in treatment markets. Every market should have a known assignment probability, and analysis should include all assigned eligible markets under an intention-to-treat approach. That prevents analysts from dropping low-delivery treatment areas after seeing sales.
Use a user-level A/B test when an audience can be individually randomized without contamination and the question concerns the product experience. Use geo lift when the intervention is market-wide or when the outcome is aggregate and media exposure cannot be individually withheld. The approaches can complement each other: geo lift establishes total incremental demand, while a site experiment tests landing-page conversion for the demand that arrives.
Worked scenario: estimating incremental streaming subscriptions
A streaming service considers a regional podcast campaign. Its media platform reports many attributed trials, but leadership wants to know whether those trials are incremental. The team chooses 24 comparable media markets, excludes markets with a scheduled price change, pairs them using twelve months of trial starts and paid retention, and randomly assigns one market in each pair to receive a six-week campaign with 25% incremental spend.
The primary outcome is net new paid subscriptions attributed to subscriber billing address, measured through 28 days after the flight; trial starts and cancellation requests are secondary. The protocol freezes other local media and tracks national brand activity. Campaign logs reveal minor spillover into two control markets. The primary analysis remains the pair-level intention-to-treat comparison; a prespecified sensitivity analysis reduces the estimated treatment intensity for those control markets.
Treated markets gain 8.4% more net new subscriptions than paired controls, with an interval consistent with 2.1% to 14.7% lift. Incremental contribution profit exceeds campaign cost under conservative retention assumptions. The team reports the total effect, not a cost per platform-attributed conversion, and runs a follow-up at a different spend level before rolling out nationally. The test supports a decision about this campaign configuration and market set, not a universal claim about podcasts.
Planning and analysis workflow
- Write a hypothesis, minimum profitable lift, target market universe, outcome definition, and campaign delivery contract.
- Audit historical outcome coverage and select markets before seeing post-period data.
- Match or stratify on pre-period level and trend; preserve a reproducible randomization record.
- Use historical simulations to set market count, duration, and precision expectations.
- Freeze competing local actions where possible and monitor spend, delivery, inventory, and cross-border behavior.
- Lock outcome data after the planned lag; analyze assigned markets with the specified pair, regression, or randomization-inference method.
- Report lift, absolute increment, uncertainty, cost, delivery fidelity, contamination, and all deviations.
Interpreting lift and return
A positive lift estimate answers a counterfactual question: outcomes were higher in markets assigned to the incremental campaign than expected under the control plan. It is not necessarily the return from all advertising, the effect of a different spend level, or the lifetime value of acquired customers. Separate the experimental effect from financial assumptions used to convert it into ROI.
Express both absolute and relative lift. Relative lift can make a small baseline look dramatic; absolute increment permits budgeting. Show an interval and compare it with the break-even effect. An interval crossing break-even may justify a cheaper confirmation test, a constrained rollout, or no action depending on downside tolerance. A non-significant result is not proof of zero increment unless the design was precise enough to exclude meaningful gains and losses.
Limitations and common mistakes
- Using platform attribution as lift: attributed conversions are not a counterfactual comparison.
- Underpowered geography count: many customer events cannot replace independent markets.
- Comparing before and after only: time trends and seasonality can mimic media impact.
- Ignoring spillover: control exposure biases contrasts toward zero or makes the estimand unclear.
- Changing spend mid-test: undocumented delivery changes make treatment intensity uninterpretable.
- Ignoring profit: revenue lift can be unprofitable after media, discounts, fulfillment, and retention are considered.
Frequently asked questions
Is geo lift the same as marketing mix modeling?
No. Marketing mix modeling is observational model-based attribution over historical variation. Geo lift is a prospective controlled experiment, though its results can calibrate a media mix model.
Can controls receive baseline advertising?
Usually yes. The contrast is often incremental spend versus the ordinary baseline plan, not advertising versus no advertising.
What is a good test duration?
Long enough for delivery, conversion lag, and normal weekly cycles, but not so long that market conditions or execution drift. Historical simulations and the outcome window should drive the choice.
Can we calculate ROI from lift?
Yes, if costs, margins, returns, and retention assumptions are explicit. The experiment estimates increment; ROI requires a transparent economic model.
Summary
A geo-lift test uses randomized treated and control markets to measure the incremental business impact of a campaign. Reliable tests specify the intervention and outcome, plan at the market level, audit delivery and spillover, and interpret lift alongside uncertainty and economics. It is a stronger answer to “what did this media cause?” than attribution counts alone.
Sources
- Vaver and Koehler, “Measuring Ad Effectiveness Using Geo Experiments”
- Kerman, Wang and Vaver, “Estimating Ad Effectiveness Using Geo Experiments”
- Lewis and Rao, “The Unfavorable Economics of Measuring the Returns to Advertising”