"Incremental" is the most load-bearing word in marketing measurement and one of the least examined. It gets used as though it were a property of a number, like being large or being negative. It is not. It is a property of a comparison, and the comparison is usually left unstated. Once you say it out loud, most arguments about effectiveness turn out to be arguments about the comparison, conducted by people who believe they are arguing about the result.
The figures and charts in this post come from a small simulation built for it, not from client work. The point of using synthetic data here is that I can plant a known truth and then show you that the observed data cannot reveal it, which is difficult to demonstrate honestly with real numbers.
The Word Hides a Subtraction
Every incrementality claim, however it is dressed up, is one subtraction.
Incremental = what happened − what would have happened anyway
The first term is data. The second term does not exist and never will.
The second term has to be constructed: from a holdout region, a matched market, a pre-period, a model, or in the worst case somebody's expectation of what a normal week looks like. So incrementality is never a measurement in the way that sales are a measurement. It is the estimated distance between a fact and a fiction, and its quality is governed almost entirely by the quality of the fiction. Two analysts with the same sales data and different baselines are not disagreeing about the past. They are disagreeing about an imaginary world, which is a much harder argument to settle and a much easier one to have unknowingly.
This is why "is it incremental?" is not answerable as asked. The answerable version has four parts, and changing any one of them changes the number, often by more than the number itself.
- Incremental to what baseline? Against no activity at all, against last year, against a cheaper alternative use of the same money?
- Incremental at what level? The SKU, the brand, the category, the retailer, the group.
- Incremental over what window? The promoted fortnight, the quarter, the year.
- Incremental for whom? Your P&L and your retailer's P&L are different questions with different answers, and both parties will quote the same lift.
Six Ways to Sell the Same Extra Units
Take a concrete case. A single grocery SKU runs a two week price promotion at one retailer, ten per cent off a product carrying a forty per cent margin. Over the two weeks it sells about 53,800 units more than its baseline, a lift of 62%. Nobody is disputing the 53,800. It is in the till data.
Here are six different worlds that produce that identical spike.
New demand. The deal brought people into the category who were not going to buy at all, or persuaded existing buyers to consume more than they otherwise would. Nothing anywhere else went down. This is the case everyone implicitly assumes and it is the rarest of the six.
Competitor switching. Buyers who would have bought a rival brand bought yours. Entirely real for you, entirely invisible in the category total. Note that this is the case where somebody else has a strong incentive to undo your gain, which makes it the least durable of the genuinely positive outcomes.
Another category. The money came out of a different aisle. The shopper had a fixed budget and a fixed occasion to fill, and your deal won it against wine, or a takeaway, or a different kind of treat entirely. Good for your brand, good for your category, worth nothing to the retailer whose total basket has not moved.
Pull-forward. The same households bought the same annual volume, earlier and more cheaply. They stocked up. Your spike is a loan against your own next two months, taken out at a ten per cent discount.
Your own second SKU. Buyers traded across your range, from the single to the multipack. The promoted line grew by very nearly what the other line lost. Genuinely incremental for the SKU, its brand manager and their bonus, and worth approximately nothing to the brand.
A rival retailer. Your own buyers bought your own product, in a different shop from usual. This one is fascinating because it is genuinely incremental for the retailer running the promotion and close to worthless for you, which is precisely why retailer media reporting and brand measurement so often disagree while both being correct.
Now the important part. In the promoted weeks, all six worlds are numerically identical. Not similar: identical. Same units, same lift, same shape, same beautiful chart for the end of quarter review. Everything that distinguishes them is either in a series belonging to somebody else, or in weeks that have not happened yet.
This figure needs JavaScript. The short version: six scenarios produce an identical promotional spike, but the share of it that survives at brand, category and retailer level ranges from 100% down to zero, and the margin impact ranges from +£14.8k to -£23.5k.
Read the margin figure at the bottom of that panel as you click through, because it is the part that decides whether any of this was worth doing. Three of the six scenarios return about £14,800 of extra gross margin over the quarter. The other three lose between £21,800 and £23,500. The promoted weeks look the same in all six cases and so does every dashboard built on them.
| Where the lift came from | This SKU | Our brand | The category | This retailer | What undoes it |
|---|---|---|---|---|---|
| New demand | Yes | Yes | Yes | Yes | Nothing, if the new buyers repeat at full price. |
| Competitor switching | Yes | Yes | No | No | Their counter-promotion, usually within a quarter. |
| Another category | Yes | Yes | Yes | No | The other category discounting back at you. |
| Pull-forward | No | No | No | No | Next month, automatically, whether you look or not. |
| Your own second SKU | Yes | No | No | No | Nothing. It was never there to begin with. |
| A rival retailer | Yes | No | No | Yes | The rival shop matching the price. |
Two things in that table are worth dwelling on. The first is that only one row is positive at every level, which should temper how casually the word gets used. The second is that the rows disagree with each other about who should be happy, and every one of those parties has access to the same till data.
Name the Loser
Which suggests a test that costs nothing and catches a surprising amount.
Every incremental unit came from somewhere. A rival's shelf, a different aisle, a different shop, your own other line, or your own next month. If you cannot name who is worse off, you have probably not found incrementality. You have found a baseline you mislabelled.
The single exception is genuinely new consumption, and it is worth being honest that in a mature category it is the least likely explanation on the list rather than the default one. Categories do grow, and advertising does grow them, but a fortnight of ten per cent off is a strange mechanism to credit with it.
The same logic scales up into a consistency check almost nobody performs. Add up what every brand in a category claims from promotion, and compare it with how much the category actually grew. If four brands each report a solid incremental gain from discounting in a category that grew by one per cent, they cannot all be right, and the interesting question becomes which of them is measuring switching and calling it growth. Each brand sees only its own model, so this contradiction can persist for years without anyone encountering it. You can usually run the check with published category data in an afternoon.
The Window Is Part of the Answer
Pull-forward deserves its own treatment, because it is the one mechanism whose evidence sits inside your own data, in your own series, where you could find it if you were looking.
The trouble is that the spike and the payback are not equally visible, even when they are the same units. A stock-up concentrates its gain into two weeks and spreads its cost over six. In our simulation the promotion adds about 53,800 units in the promoted fortnight and gives 85% of them back afterwards. The spike is a 62% deviation arriving all at once, and it is unmissable. The payback starts around a third below baseline and fades to nothing over the following month and a half, so it never presents itself as an event. It presents itself as a soft patch, which is the sort of thing that gets attributed to the weather, or a competitor, or nothing at all.
This figure needs JavaScript. The short version: measured over two weeks both scenarios return the same +£14.8k of margin. Measured over thirteen, one is still +£14.8k and the other is -£21.8k.
Read across that chart and the promotion is a success, a wash, or a serious mistake depending only on where you stop. And notice what sets the window in most organisations: the reporting cadence. A four week post-campaign readout exists because the meeting is monthly, not because four weeks bears any relationship to how often somebody buys washing powder.
If your evaluation window is shorter than the purchase cycle, you have guaranteed yourself a flattering answer before you have looked at any data.
What This Does to Your Model
All of which lands on the model, and on five specific decisions that people tend to treat as technical detail rather than as the choice of what question is being answered.
A model can only find mechanisms you gave it a way to express
This is the most important one and the least discussed. If your specification contains no term capable of representing a post-promotion dip, the model cannot report pull-forward. It will not warn you. It will fit those quiet weeks with whatever it does have available: a slightly lower intercept, a seasonality wiggle, a competitor coefficient doing unexpected work. The output will look complete. The residuals may look fine.
Absence of a mechanism in your specification is not evidence of absence in the world, but in a results deck the two are indistinguishable. Anything you cannot express, you will implicitly estimate as zero, and you will report that zero with a confidence interval around it.
The baseline is the estimand, and you chose it
Whatever you allow to soak up variance determines how much is left for marketing to explain. A flexible weekly baseline will cheerfully absorb your promotional spike as "an unusual week". A rigid one hands the whole spike to whichever variable happens to be switched on. Both will fit the data acceptably. Neither is neutral, and the difference between them is not a technicality, it is most of the answer.
So when somebody presents an ROI, the most informative question is rarely about the media coefficient. It is: what does your baseline do during the activity, and why is it allowed to do that?
Aggregation sets a ceiling on the questions you can answer
A brand-level model cannot detect cannibalisation within the brand, because the two SKUs have already been added together before the model sees them. A category-level model cannot detect switching. A national model cannot detect one retailer's gain at another's expense. Sum first and the answer is not merely hard to obtain, it is unavailable at any price.
I have written about an expensive version of this, where the single most valuable variable in a dataset was aggregated away before modelling began and the finding sat undiscovered in a column nobody had asked for. The lesson generalises: the unit of analysis is not a convenience, it is a decision about which mechanisms remain visible.
Identification, not machinery
If promotions always run in the same weeks of the year, promotion and seasonality are collinear, and no amount of Bayesian sophistication will separate them. The prior will decide the answer, and in the output it will look exactly like evidence. Before asking how a model was fitted, ask where the identifying variation comes from, because if the honest answer is "nowhere", the rest is decoration.
Your priors decide what answers are reachable
A strictly positive prior on a promotional coefficient cannot produce a value-destroying promotion. If the model is structurally incapable of expressing "this discount destroyed margin", then its failure to say so carries no information at all. The support of a prior is a claim about which worlds are possible, and it is worth making that claim deliberately rather than by inheriting somebody's default.
Invert. Always Invert.
Charlie Munger's favourite mental tool was borrowed from the mathematician Carl Jacobi, who used to tell students man muss immer umkehren: one must always invert. Munger's own version was that it is not enough to think about a difficult problem one way, you have to think about it forwards and backwards. Incrementality is exactly the kind of question that yields more to the backwards pass than the forwards one.
Four inversions, in rough order of how much time they save.
Instead of "what did the promotion add?", ask "what would have happened if we had not run it?" These sound like the same question and they send you to different places. The first sends you to the promoted weeks, which is where the flattering evidence lives. The second sends you to the weeks either side, the competitor, your other SKUs and your other retailers, which is where the disconfirming evidence lives.
Instead of "what is the ROI?", ask "what would have to be true for this ROI to be wrong?" Then go and measure that thing. "This is a 1.9 return" invites agreement or disagreement and settles nothing. "This is a 1.9 return unless roughly a third of the volume was pulled forward, and here is what the following six weeks look like" is a specific, checkable claim. It converts an argument about a number into a task.
Instead of "how much did we gain?", ask "who lost?" As above. It is astonishing how often nobody can answer, and how quickly that reframes the conversation.
Instead of estimating the effect, estimate the break-even. This one is underused and it is the strongest of the four. Rather than asking what the incremental volume was, ask how much of the spike would have to be genuinely new for the promotion merely to pay for itself. In our example the arithmetic is unavoidable: a ten per cent price cut on a forty per cent margin means every unit in the promoted weeks earns a quarter less, including all the units that were going to sell anyway. Working backwards, 54% of that 53,800 unit spike has to be genuinely new before the promotion breaks even at brand level.
That is a far more robust number than any ROI estimate, because it depends only on price, margin and baseline volume, all of which you know, and not at all on a counterfactual you have to model. It also turns the whole debate into a single question a commercial team can actually reason about: is it plausible that more than half of this was new? Often, stated that plainly, everyone in the room already knows the answer.
Stress Testing, Because One Run Is Not a Measurement
A model that has been run once, on one specification, is not a measurement instrument. It is an opinion with a confidence interval attached, and the interval describes uncertainty about the coefficients rather than uncertainty about whether the model is the right one. The second kind is usually much larger.
The most valuable test is also the one most often skipped: plant a truth and see whether your pipeline can recover it. Simulate data where you know the answer, run your actual modelling process over it, and compare. Do it with the awkward cases rather than the easy ones: a promotion that is 60% pulled forward, a competitor who responds two weeks later, activity that always coincides with a seasonal peak, a cannibalising second SKU. If your process cannot recover an effect you planted yourself, you have no basis for trusting it on data where nobody knows the truth. This costs a day and it is the closest thing to a calibration certificate that our field has.
The second test is a specification sweep. Take each defensible choice, swing it to the other defensible setting, and record what happens to the answer.
This figure needs JavaScript. The short version: six defensible modelling choices, each swung between two reasonable settings, move the reported return on the same promotion from 0.42 to 1.94. Most of the ranges straddle break-even.
The uncomfortable implication of a chart like that is not that modelling is hopeless. It is that the honest deliverable is sometimes a range plus a recommendation. If a defensible choice can move the answer across break-even, then the correct output is "we cannot yet tell whether this promotion pays, and here is the one measurement that would settle it", which is very often an experiment rather than more modelling.
A fuller battery, worth running before anybody makes a decision on the output.
| Test | What it catches | Red flag |
|---|---|---|
| Recover a planted truth | A pipeline that cannot find real effects, or invents absent ones | The estimate misses the planted value by more than its own stated uncertainty |
| Specification sweep | Answers that depend on arbitrary choices rather than on data | The range crosses break-even, and only one end of it gets presented |
| Time holdout | Overfitting, and baselines that are really absorbing noise | Out-of-sample error far exceeds in-sample residuals |
| Extreme inputs | Functional forms that misbehave outside observed ranges | Ten times the spend implies ten times the sales, or zero spend still implies 95% of revenue |
| Placebo and shuffled dates | A model that finds structure in noise | A clear effect on a period when nothing was running |
| Negative control outcome | Confounding by seasonality, distribution or price | Your activity appears to move a product it could not possibly affect |
| Reconcile with an experiment | Everything above, at once | Model and test disagree beyond their intervals and the disagreement is not investigated |
Notice how much of that battery is designed to make the model fail rather than to confirm it. That is the inversion again, applied to your own work. The forwards question is "does my model fit?" The backwards question is "what would I expect to see if my model were wrong, and do I see it?" Only the second one can change your mind.
Incremental, and Yet Not New
One more case, because it shows that the honest answer is often "yes, and here is exactly what kind".
In the £14,999 story, a lender changed one number in a pricing table and recovered around 215 loan applications a week. That was unambiguously incremental to the lender: real applications, real advances, real net present value, and it was the largest single reason a target was met. It also created essentially no demand. The people who wanted to borrow £15,000 wanted to borrow it the whole time and were going to eleven competitors who looked cheaper on the comparison page.
Put through the four questions, that result is easy to state precisely. Incremental against a baseline of the eighteen months in which the boundary was mispriced. Incremental at brand level, close to zero at category level. Durable until a competitor repriced. And the losers were nameable to within a shortlist of eleven. That is a well-formed incrementality claim, and notice that it is a stronger claim for being narrower, not a weaker one.
Harvesting demand that already exists and creating demand that does not are both worth money. They are not worth the same money, they do not have the same ceiling, and they should not be planned, forecast or rewarded as though they were the same thing.
Four Questions to Staple to Any Number
None of this requires new tooling. It requires refusing to accept the word on its own.
- Incremental to what baseline, and what does that baseline do during the activity?
- Incremental at what level: SKU, brand, category, retailer, group?
- Incremental over what window, and how does that window compare with the purchase cycle?
- Who is worse off, and would they recognise the description?
The word promises a fact and delivers a comparison. Getting into the habit of asking which comparison is, in my experience, worth more than any amount of additional modelling sophistication applied to the wrong one.
Want a second opinion on an incrementality number?
We audit models, build in-house measurement capability, and spend a good deal of our time helping teams work out what their numbers are actually a comparison against. Happy to look at yours.
Book a Discovery Call