“The board wants proof that marketing works”

Nothing here produces proof. What they produce is evidence of different strengths and costs — an experiment is the strongest and the slowest, a model is the broadest and the most assumption-laden. Say which you are offering.

10 tools across 2 of the 9 questions

8.5 Evidence tier: Law

Incrementality / Geo-Lift Testing

A randomised experiment that withholds spend from matched markets or audiences in order to measure what that spend actually caused.

Breaks when Needs sample, time and patience. The statistical power to detect a small effect is absent from most budgets, so “no effect” usually means “we could not measure it”.

Allocate budget
8.4 Evidence tier: Frame

Marketing Mix Modeling

A regression on aggregate historical data that estimates each marketing input's contribution to sales.

Breaks when Correlational. Uncalibrated by experiment it moves budget systematically to the wrong channel — and whoever builds the model can largely choose the answer.

Allocate budget
9.7 Evidence tier: Frame

Test & Roll

An experiment design that sizes a test to maximise total profit across the test and the rollout, instead of sizing it to reach statistical significance.

Breaks when It optimises the profit of one decision, so it deliberately accepts a higher error rate than a scientific test would — it will sometimes roll out the worse arm, by design, because the expected cost of that is lower than the cost of the larger test. That makes it the wrong instrument when the result has to generalise, feed a model, be reused across markets, or be defended to someone outside the team. It also needs a prior on the effect size, and a badly chosen prior moves the recommended sample size a long way.

Measure
9.6 Evidence tier: Frame

Share of Search

A brand's share of category-related search volume, used as a fast and free leading indicator of market share.

Breaks when It is a proxy, and proxies drift from what they proxy. Search volume rises with newsworthiness as well as with demand, so a crisis and a successful campaign look the same. In categories where buying does not involve search, or where the brand name is also an ordinary word, the signal is mostly noise. And because it is cheap to move — brand-name bidding, PR spikes — it degrades the moment it becomes a target, which is the standard Goodhart path.

Measure
9.2 Evidence tier: Frame

Brand Tracking

Repeated survey measurement of a brand's presence in memory — recall, consideration, asset attribution — reported as a time series.

Breaks when Question wording determines the answer; mix prompted and unprompted recall and the series becomes meaningless. Change methodology and the entire history is void.

Measure
9.3 Evidence tier: Folklore

Attribution Models

Rules that assign credit for a conversion across the touchpoints that preceded it.

Breaks when Observational, not causal. Benchmarked against experiments it systematically inflates cheap and search channels. It cannot substitute for an incrementality test, and it constantly is.

Measure
9.5 Evidence tier: Law

Goodhart's Law

The principle that an observed statistical regularity collapses once it is used as a target for control.

Breaks when It does not break. What breaks is the measurement system that forgot it.

Measure
9.8 Evidence tier: Law

Split-Run Testing

Running two versions of the same advertisement to randomly divided halves of one audience and keeping the version that produces more response.

Breaks when Most tests are underpowered for the difference they are looking for, so a large share of declared winners are noise that will not repeat. The design also only compares the variants you thought of: it optimises inside a set and never reports that the set was wrong, which is how a headline test returns a clean 12% lift on a page selling the wrong thing. And it measures response, not profit — Hopkins' keyed coupon counted replies and the modern equivalent counts clicks.

Measure
8.6 Evidence tier: Law

Advertising Elasticity

The percentage change in sales produced by a 1% change in advertising spend, generalised across studies into a single magnitude.

Breaks when It is an average, and it has been falling. Elasticity is higher for durables than non-durables, higher early in the life cycle than at maturity, and higher when advertising is measured in gross rating points than in money — so the headline figure describes no actual brand. Used as a planning input it also silently assumes average-quality advertising, which is the variable most under a team's control and precisely the one the meta-analysis averages away.

Allocate budget
9.9 Evidence tier: Frame

Customer-Based Brand Equity

Brand equity defined as the difference in how a customer responds to marketing for a branded product versus the identical product unbranded, located in brand knowledge held in memory.

Breaks when It never specifies the weights. Nothing in the framework says how much awareness is worth against how much favourability, so every tracker built on it makes those trade-offs implicitly and reports the result as a score. Defined as a differential response, it is also close to unfalsifiable in practice: any gap between a brand and an unbranded equivalent counts as evidence for it, and the framework names no case in which the differential should be zero.

Measure