Skip to content

Guide · 9 min read

Incrementality TestingMeasuring what would not have happened

Diagrams
02
Tools
04
Sections
08

The short answer

Incrementality testing measures the extra sales, leads or other outcomes that marketing causes, by comparing a group exposed to it with a similar group that is not. Common designs are audience holdouts, geographic experiments and platform lift studies. It is the most direct way to separate what marketing caused from what would have happened anyway.

The counterfactual question

Every marketing result has an invisible twin: what would have happened without it. A retargeting campaign that reaches a hundred cart abandoners who then buy looks spectacular, until you ask how many of them would have come back and bought anyway. Incrementality is the difference between the two.

Attribution cannot answer this question because it only sees the people who were exposed. Incrementality testing creates the missing comparison deliberately: a control group that does not receive the marketing, chosen so that it is otherwise as similar as possible to the group that does.

The most expensive marketing is the kind that takes credit for sales you already had.

The three main designs

Compare scenarios

Choosing a test design

A random share of the target audience is withheld from the campaign. Outcomes in the exposed and held-out groups are compared.

  • Best for CRM, email, retargeting and app messaging
  • Randomisation makes it clean
  • Needs reliable identity and enough volume

Our default recommendation is the geo experiment for any channel with meaningful spend, because it works without user-level tracking and can be run on search, social, video, outdoor or radio alike. Audience holdouts are the gold standard where you control the audience directly, such as email or retention marketing.

Designing a fair test

Fig. 01 · Process

Seven steps to a test you can trust

Most failed tests fail at steps two and three, before any money is spent.

Write the decision first. 'If paid social adds less than X in incremental gross profit per rupee, we cut it by a third' is a test brief. 'Let us see what happens' is not. Agreeing the action in advance prevents a disappointing result from being explained away.

Measure in your own systems. The outcome should come from your order system, CRM or finance data, not from the platform being tested. A platform measuring its own lift has an obvious conflict, even when its methods are sound.

Size the test. Small tests produce noisy results that cannot distinguish a modest effect from zero. Before launching, estimate the smallest effect that would change your decision and check, with an analyst or a power calculator, that your volumes and duration can detect it. If they cannot, enlarge the test or choose a different question.

Geo experiments in practice

A geo test divides your market into regions (states, cities, postal clusters or the regional targeting units your ad platforms support) and assigns some to a changed treatment. The art lies in matching. Test and control regions should have behaved similarly in the past, so that any divergence during the test can be attributed to the change rather than to local quirks.

In a country as varied as India, matching deserves real care. Metro and tier-2 markets differ in device mix, payment habits, delivery coverage and festive calendars; regional festivals fall at different times in different states. Match on historical sales patterns rather than intuition, exclude regions with known disruptions such as a store opening, and prefer several smaller regions per group over one large one.

Analysis typically uses the control regions to forecast what the test regions would have done without the change, then measures the gap. Synthetic-control methods and open-source geo-experiment libraries make this more rigorous; an analyst with experience of them is worth involving for any test that will move significant budget.

Reading the result

The basic arithmetic is simple: compare the conversion rate (or revenue per person, or sales per region) in the test group with the control group. The difference is the lift. Multiply the lift by the size of the exposed group to estimate incremental outcomes, and divide spend by those outcomes to find the true cost per incremental result.

Calculator

Holdout test calculator

Illustration only. Enter the results of an audience holdout test. Outputs ignore statistical significance, which you should check separately.

Exposed conversion rate

2%

Rate among people who could see the campaign.

= testconv / testn

Holdout conversion rate

1.7%

What would have happened anyway, estimated.

= ctrlconv / ctrln

Incremental conversions

270

Conversions the campaign caused, estimated.

= (testconv / testn - ctrlconv / ctrln) * testn

Share of conversions that were incremental

15%

Compare with what attribution credited.

= (testconv / testn - ctrlconv / ctrln) * testn / testconv

Cost per incremental conversion

₹1,852

The number that should drive budget.

= spend / ((testconv / testn - ctrlconv / ctrln) * testn)

Defaults are illustrations. Use your own numbers. Nothing you enter leaves this page.

Always report a range, not just a point estimate. A lift of a given size with a wide confidence interval that includes zero is not evidence the channel works; nor is it evidence that it does not. It is evidence the test was too small. Say so plainly, and resist the temptation to re-slice the data until some segment looks significant.

Using results: calibration, not trophies

A test result is most valuable when it changes how you read everyday data. If a holdout shows that only part of the conversions attributed to retargeting were incremental, apply that ratio as a correction factor to retargeting's attributed results until the next test. Do the same for each major channel. Your attribution reports become calibrated estimates rather than raw claims.

Results also feed marketing mix modelling as anchors and inform CAC calculations: incremental CAC is the honest version of the number boards care about.

Fig. 02 · Matrix

Where to test first

LargeShare of budgetSmall
LowConfidence in current measurement →High
Prioritise channels where spend is large and your confidence in their effect is low.

Common mistakes

Myth vs reality

Incrementality misconceptions

Other frequent errors: changing creative or budget mid-test, choosing control regions that differ in obvious ways (a metro against a tier-2 city), testing during a festive peak without accounting for it, and reading results daily and stopping when they look good. Each introduces bias that can overturn the conclusion.

Building a testing habit

The organisations that benefit most treat testing as a calendar, not an event. They keep a backlog of questions ranked by budget at stake and uncertainty, run a test each quarter on the top item, and publish results internally whether they flatter anyone or not. Over a couple of years this builds a library of evidence that no attribution tool can match.

Checklist

0/7

Test brief template

Key takeaways

  1. 01Incrementality is the difference between outcomes with marketing and outcomes without it.
  2. 02Geo experiments work for almost any channel; audience holdouts are ideal where you control the audience.
  3. 03Agree the decision and the success threshold before the test, and measure outcomes in your own systems.
  4. 04Report results as ranges; an inconclusive test usually means it was too small, not that nothing happened.
  5. 05Use results to calibrate attribution and mix models, and re-test major channels regularly.

Frequently asked

What is the difference between incrementality and ROAS?
ROAS is revenue attributed to advertising divided by spend, using an attribution model that credits touchpoints. Incremental ROAS uses only the revenue the advertising actually caused, as measured by a test with a control group. Incremental ROAS is usually lower, sometimes much lower, especially for retargeting and branded search.
How long should an incrementality test run?
Long enough to collect sufficient conversions to detect the smallest effect that matters, and to cover at least one full purchase cycle and weekly patterns. That might be a few weeks for high-volume ecommerce or several months for B2B. Estimate the required duration before launching, not during the test.
How big should a holdout group be?
Large enough to measure the baseline conversion rate precisely, which depends on your volumes and the effect size you want to detect. Smaller holdouts cost less in lost sales but produce noisier results. A power calculation balances the two; avoid choosing a size by habit.
Can I run incrementality tests without user-level tracking?
Yes. Geo experiments compare regions rather than individuals, using aggregate sales or leads by region from your own systems. They are well suited to a privacy-constrained world and to offline media, although they need careful region matching and analysis.
Are platform conversion lift studies reliable?
They often use sound randomised methods and are a useful, convenient source of evidence. Their limitation is that the platform designs, runs and reports the test of its own effectiveness, and eligibility or methods may change. Treat them as one input and confirm decisive findings with an independent design.

Published by Fabulous.Media, a network of specialist marketing agencies. Updated 9 October 2026. Platform features change often; check current official documentation before acting on platform-specific detail.

Read next

Prefer a specialist to do this with you? The network has a house for every discipline in this library.

Request an Introduction