Design·Glossary term

Geo Experiment

Geo Experiment A/B testing Reference guide

Geo Experiment is a concept used in experiment design & methodology.

Quick definition: A geo experiment estimates an intervention’s causal effect by assigning geographic areas, rather than individual people, to treatment or control and comparing their outcomes over a defined period.

What is a geo experiment?

A geo experiment, also called a geographic randomized experiment or geo test, is a controlled study in which the unit of assignment is a market: a country, region, city, postal area, media market, store catchment, or other defensible geography. Treated areas receive an intervention such as additional advertising, a different promotion, a new delivery policy, or a product launch; matched control areas do not. The resulting comparison estimates incremental impact within the tested places and period.

It is useful when people cannot reliably be randomized one by one. Marketing impressions can spill across devices, offline media reaches a whole local audience, prices are set by store, and delivery operations are often regional. A person-level A/B test may appear feasible but violate stable treatment exposure when household members, stores, or local awareness affect each other. Randomizing at the geographic level makes the intervention and analysis unit better aligned.

A geo experiment is not merely comparing sales in two cities. Causality comes from planned assignment, credible control markets, concurrent measurement, and a predeclared analysis. A city chosen because it had weak sales after a campaign is an observational case study, not a randomized geo test.

Design methodology

First define the intervention, estimand, and market universe. The estimand might be incremental weekly orders per eligible household caused by a two-week video campaign, or incremental gross profit caused by delivery-fee removal. Define treatment exposure precisely: spend level, channels, start and stop times, creative, geographic targeting, and any planned exceptions. Define the outcome source, attribution window, and whether the outcome is store sales, online orders, revenue, profit, registrations, or a blended metric.

Markets are commonly paired or stratified using pre-period outcome level, trend, population, seasonality, channel availability, and operational characteristics. Within pairs or strata, randomly assign one market to treatment and one to control. Matching improves precision but does not replace randomization. With many markets, constrained randomization can select allocations balanced on key baseline variables while retaining a documented random mechanism.

Design elementWhy it mattersEvidence to retain
Market definitionSets the unit receiving the intervention and avoids overlap.Boundary map, exclusions, and customer-to-market rule.
Pre-period matchingReduces avoidable baseline imbalance.Historical levels, trends, and matching algorithm.
Random assignmentSupports a causal treated-versus-control contrast.Seed, allocation file, and assignment timestamp.
Concurrent measurementProtects against calendar-wide shocks.Locked outcome extracts and common calendar windows.

Sample size is often the hardest constraint. The independent observations are markets, not customers. A million transactions from four treated cities do not supply the same experimental information as a million independently randomized users. Plan from the number of areas, expected between-market variation, pre-period correlation, expected effect, and possible interference. When there are few clusters, rely on randomization inference, permutation tests, or carefully justified small-sample methods rather than a large-sample individual-level standard error.

Assumptions and validity conditions

Random assignment makes treated and control markets comparable in expectation, but execution can undo that advantage. Geographic targeting must match assignment; campaign delivery needs to be audited for leakage, frequency, audience composition, and spend. A treatment market that receives half the intended media and a control market that receives spillover is a different intervention from the protocol.

The central challenge is interference. Customers can commute, shop across borders, see national media, share referral codes, or order from a nearby store. Suppliers and competitors may react to a local campaign. Define buffer zones or exclude border areas when necessary, and measure cross-market purchasing. A geo estimate can remain valuable under limited spillover, but its estimand becomes the effect of the campaign as actually delivered in a connected market system, not a perfectly isolated local treatment.

Outcome measurement must use stable identity and aggregation rules. Do not remove low-performing markets after assignment or change a revenue definition because a campaign affects discounts. Check for pre-period trend balance, major local shocks, stockouts, tracking changes, holidays, and concurrent promotions. These checks diagnose credibility; they should not become a license to choose the analysis that produces the preferred answer.

Geo experiments and A/B testing

A geo test applies the same controlled-experiment logic as an individual A/B test, but with clusters. Treatment and control must be concurrent, eligibility must be clear, and the primary outcome and decision rule should be set before observing the post-period result. Inference must respect the randomization unit. An analysis that treats every transaction as independent can report implausibly narrow intervals because transactions within a city share weather, local demand, inventory, and media exposure.

Use an individual-level A/B test when exposure can be randomized and isolated per user. Prefer a geo experiment for market-level ads, store operations, region-wide policies, or when individual targeting creates unacceptable contamination. In some programs, a geo experiment validates an aggregate marketing effect while user-level product tests optimize the experience after acquisition.

Worked scenario: measuring incremental local search advertising

A retailer wants to know whether increasing local-search advertising produces incremental store profit, not just attributed clicks. It selects 30 media markets with stable point-of-sale coverage and excludes markets with remodels or persistent inventory problems. Markets are grouped into 15 pairs using pre-period weekly profit, online demand, population, and historical response to promotions. One market in each pair is randomly assigned to 40% higher local-search spend for six weeks.

The primary outcome is weekly contribution profit per market; online orders delivered across a market border are assigned by delivery address. The protocol freezes national advertising and prohibits local coupons outside the treatment campaign. A two-week washout is planned because some search response may persist. During the test, the team verifies actual spend, impression geography, inventory availability, and cross-border sales. One control market receives accidental treatment spend for two days; that deviation is documented and handled under the prespecified sensitivity analysis rather than silently dropped.

After the campaign, treated markets show a 5.2% higher average profit change than their paired controls. The interval is wide because there are only 15 pairs, but it excludes the minimum effect needed to cover spend. The decision report gives the intent-to-treat estimate, actual delivery diagnostics, paired-market results, and sensitivity to the contaminated market. The company funds a broader rollout with ongoing holdout monitoring rather than claiming that every city will gain exactly 5.2%.

Analysis workflow

  1. State the market-level business decision and define treatment, control, eligibility, outcome, and required observation window.
  2. Audit historical data and remove or flag markets with unusable measurement before randomization.
  3. Match or stratify markets on pre-period level and trend, then create a reproducible random allocation.
  4. Estimate the needed number of markets through simulation using historical covariance and plausible spillover.
  5. Lock campaign delivery and operational policies; monitor fidelity without selecting results.
  6. Analyze at market or matched-pair level using the prespecified model, with cluster-appropriate uncertainty and diagnostics.
  7. Report absolute outcomes, incremental effect, uncertainty, deviations, spillover evidence, and the rollout decision.

Interpreting a geo-test result

The primary result estimates the average effect of assignment to the geographic intervention in the experimental setting. It may be an intention-to-treat effect: the effect of offering the planned campaign, including ordinary delivery imperfections. That is often the most useful business quantity. Do not relabel it as the effect per ad impression or per exposed person without a separate, well-supported analysis.

A statistically uncertain estimate is not automatically evidence that the campaign has no value. Compare the interval with the minimum profitable effect and the cost of more experimentation. Conversely, a precise positive estimate does not establish that the result will transfer to untested markets, another season, a different spend level, or a national rollout where media auction prices and competition change.

Limitations and common mistakes

  • Too few markets: transaction volume cannot compensate fully for a tiny number of independent geographic units.
  • Spillover blindness: untreated customers may see treatment media or buy in treatment stores.
  • Post hoc matching: selecting controls after outcomes are known invites cherry-picking.
  • National confounding: a one-period before-and-after comparison cannot isolate a campaign from broad trends.
  • Delivery drift: unequal spend, targeting, inventory, or promotions means the tested treatment was not consistent.
  • Individual-level inference: ignoring clustering makes uncertainty artificially small.

Frequently asked questions

How many geographies are enough?

There is no universal minimum. Plan from effect size, market variance, matching quality, and required precision, but recognize that more independent markets are usually more valuable than more transactions in the same markets.

Can a city be both treatment and control?

Not for one indivisible city-level intervention. It may be possible to randomize smaller non-overlapping zones, provided targeting, customer movement, and operations can preserve those boundaries.

Does matching prove causality?

No. Matching improves efficiency. Random assignment and faithful execution provide the causal foundation.

Should we exclude a market with poor performance?

Only under an exclusion rule written before results are examined. Report every exclusion and sensitivity analysis.

Summary

A geo experiment randomizes markets to estimate the incremental effect of an intervention that cannot be cleanly assigned to individuals. Strong designs define the geographic unit, use reproducible assignment and concurrent controls, respect the small number of independent areas, and audit spillover and delivery. Its conclusion is local and conditional: useful for the tested intervention and markets, not a promise that all future geographies will respond identically.

Sources