Attribution tells you which campaign received credit. Incrementality asks the harder question: what would have happened if the campaign had not run?
That counterfactual cannot be observed directly. It has to be estimated from a control group or comparison market. The quality of the answer depends less on choosing the most sophisticated method and more on matching the test design to the decision the business needs to make.
Start with the decision, not the method
“Is this campaign incremental?” is too broad to design a useful test. A growth team may actually need to decide whether to increase budget, enter a market, keep a retargeting tactic, change a bid strategy or compare the lift of two supply packages.
Write the decision before the experiment:
If measured lift clears [threshold], we will [specific action].This statement forces clarity on the outcome, population, time window and minimum useful effect. It also prevents a common failure mode: running a test, seeing a noisy result and inventing the decision rule afterward.
The three core methods
User-level holdouts
Eligible users are randomly assigned to treatment and control. The treatment group can receive the campaign; the control group is intentionally withheld. Randomization makes the groups comparable in expectation, so the difference in outcomes estimates incremental lift.
User holdouts are intuitive and often precise when identity and eligibility can be managed consistently. They become harder when users move across devices, when exposure happens through several uncoordinated channels or when withholding at the user level is not operationally possible.
Ghost-bid experiments
A ghost-bid design creates a control from auction opportunities the campaign was eligible to win, then withholds delivery for the control assignment. Because treatment and control originate from comparable auction opportunities, the method can reduce bias created by comparing reached users with everyone else.
The design requires platform support and careful handling of auction eligibility, assignment and downstream matching. It is useful when the question sits close to the media-buying mechanism itself.
Geo experiments
Geographic markets are assigned different media pressure, and outcomes are compared against matched or modeled controls. Geo tests can capture cross-device and cross-channel effects without requiring user-level identity.
The tradeoff is sample size: there are fewer geographies than users, markets differ, and spillover can weaken the contrast. Strong pre-period matching and stable market definitions matter.
A practical way to choose
Can eligible users be assigned and withheld reliably?
User holdoutIs the decision specifically about auction-driven delivery?
Ghost bidDo you need cross-device or market-level lift?
Geo testThese are not absolute rules. A team may use more than one method over time: user holdouts for always-on monitoring, then a geo experiment to validate total market lift. The methods should triangulate the same business reality, even though the measured populations differ.
Five design choices that protect the result
- Pre-register the outcome. Choose the primary success event and analysis window before looking at results.
- Estimate power. Use baseline conversion, expected effect and available population to determine whether the test can answer the question.
- Protect assignment. Keep users or markets in their assigned group and measure contamination where possible.
- Hold other changes steady. Promotions, product releases and channel shifts can overwhelm the effect being tested.
- Use one decision window. Avoid checking repeatedly and stopping only when the result looks favorable.
The control group has an opportunity cost because some media is withheld. That cost should be sized deliberately. A tiny holdout may produce an inconclusive answer; an unnecessarily large one gives up more short-term outcomes than the learning requires.
Read lift, uncertainty and economics together
A result should report more than a single lift percentage. At minimum, pair the estimated effect with its uncertainty, the size of the exposed population and the economics of the decision.
An inconclusive result does not necessarily mean zero lift. It may mean the test was underpowered, the outcome was too rare, the run was too short or the groups were contaminated. Diagnose the design before turning statistical uncertainty into a business conclusion.
Likewise, statistical evidence is not the same as economic importance. A precisely measured small effect can still be unprofitable; a promising larger effect may justify another test before scale.
Incrementality launch checklist
- Write the budget or product decision the test will inform.
- Choose the unit of assignment: user, auction opportunity or geography.
- Define the primary outcome, window and minimum useful effect.
- Check baseline balance and document possible contamination.
- Agree on the action for positive, negative and inconclusive results.
Incrementality is not a ceremonial validation step after attribution. It is a way to make budget decisions against a credible alternative reality. Choose the design that matches your control surface, protect the experiment and read the result in business terms.
