Every marketing blog lists the tools. Two columns are almost never published: how much evidence actually sits under each one, and the condition under which it stops working. This index publishes both.
tools
The set of situations, needs and moments in which a buyer thinks of a category — the cues a brand must be linked to in memory to be considered at all.
Breaks when Useless when the category does not yet exist — a genuinely new product has no moment to be recalled in, and must build one before this tool has anything to work on.
The empirical regularity that brands with smaller market share have both fewer buyers and slightly lower loyalty among the buyers they have.
Breaks when Weakens where switching is contractually locked (telco, insurance) and in heavy B2B where the buyer universe is structurally small.
A stochastic model that predicts, from category structure alone, every brand's expected penetration, purchase frequency and buyer overlap.
Breaks when Assumptions fail with few players, high involvement, or one-off purchases. It also tells you nothing about why a deviation exists.
A method that defines demand by the progress a customer is trying to make, rather than by product category or customer profile.
Breaks when In low-involvement, habitual purchases it degenerates into inventing a job that was never there. Not every purchase has a purpose behind it.
A two-factor account of brand growth: the probability of being noticed or recalled in a buying situation, and the ease of buying once you are.
Breaks when Where distribution is fixed and single-channel, the second term carries no information and everything collapses into the first.
An account of how an innovation spreads through a population over time, and of the innovation attributes that set the rate.
Breaks when Strong at explaining backwards, weak at predicting forwards. You only know who the early adopters were afterwards.
The claim that a discontinuity in buying logic separates early adopters from the early majority, crossed by dominating one narrow beachhead first.
Breaks when No empirical support for the chasm itself. Outside B2B technology the mechanism has little to stand on — the focus advice is good, the theory under it is not.
A three-layer market sizing convention: the total market, the portion a given business model can serve, and the share it could realistically win.
Breaks when Produces no decision. The number is almost always reverse-engineered from the desired conclusion, and it never identifies which assumption is the fragile one.
Dividing a market into groups expected to behave differently, then choosing which of those groups to serve.
Breaks when It breaks as a description of who buys. Competing brands within a category are repeatedly measured as having near-identical buyer profiles: the segments differ in size, not in kind. A segmentation that assigns each brand its own distinct demographic or attitudinal buyer is describing a market that measurement does not find. It survives as a targeting and media convenience, not as a theory of demand.
A composite fictional buyer — name, age, role, frustrations — built so a team has a concrete person to design and write for.
Breaks when It breaks the moment the fiction is treated as evidence. Personas are typically built from a handful of interviews and then used to exclude — this campaign is not for her, that channel is not where he is. Because competing brands largely share the same buyers, a persona used to exclude is excluding people who already buy from you. The failure is silent: nobody audits the customers a persona quietly wrote off.
The proposition that at any given moment roughly 95% of business buyers are not in the market, so most advertising has to work by being remembered later rather than by converting now.
Breaks when The 95/5 split is a derived estimate, not a measurement. It comes from average purchase-cycle length: divide the share of buyers who transact in a year by the length of the cycle and you get the in-market proportion. In categories with short cycles, high growth, or frequent re-purchase, the in-market share is materially higher and the rule under-invests in capture. It is a reminder of proportion, not a constant.
The four decision areas a marketing plan is expected to cover — product, price, place, promotion.
Breaks when It breaks the moment it is used as a strategy tool. It is a list of levers with no theory of which lever matters, in what order, or by how much — four boxes of equal visual weight for decisions of wildly unequal consequence. It is also entirely firm-side: all four Ps are things the company does, and the customer appears nowhere. The long parade of proposed extensions — 7 Ps, 4 Cs, SAVE — are symptoms of a frame that organises without prioritising.
A survey rule of thumb: ask users how they would feel if they could no longer use the product, and treat 40% answering “very disappointed” as the threshold for fit.
Breaks when The 40% threshold has no published derivation — it is a remembered benchmark from a small set of startups, not an estimate with an interval around it. Worse, the survey only reaches people still using the product, so it measures the enthusiasm of survivors and is structurally silent on everyone who already left. A product with a small, devoted, unscalable audience passes it comfortably; that is the exact failure mode it is meant to catch.
The decision about how many brands a company runs and how they relate — one master brand, a house of separate brands, or a structure between the two.
Breaks when It gives no rule for where the line falls. The frame will describe any portfolio you already have and endorse almost any addition to it, because the case for a master brand (efficiency) and the case for a separate brand (focus) are both always available. The evidence on extensions is conditional — similarity between the extension and the parent, and the strength of the parent, both matter — so a frame that does not carry those conditions is a vocabulary, not a decision rule.
The practice of claiming a single defensible place in the prospect's mind relative to competitors, rather than competing on product attributes.
Breaks when Fails when you do not control the category definition, or when the category does not exist yet. And unless the mental position is actually measured, it decays into an argument about taglines.
A five-step procedure that derives a market category from competitive alternatives, unique attributes, the value they produce, and the segment that cares.
Breaks when Powerful only where you can choose the category. In regulated or mature markets step five is locked, and the canvas quietly becomes a four-step exercise.
The strategy of creating and naming a new market category rather than competing inside an existing one, on the premise that whoever defines a category captures most of its value.
Breaks when Heavy survivorship bias. The often-quoted 76% of category market cap comes from the authors' own unpublished analysis with no disclosed sample, method or period — and the companies that tried to create a category and failed are not in the book.
The distinction between attributes on which a brand must merely match competitors and the attribute on which it must differ.
Breaks when Let the parity list grow and the product becomes a me-too. Matching is not free — most budgets disappear into it without anyone noticing.
A spatial representation of how customers perceive competing brands, derived from survey data, used to locate unoccupied positions.
Breaks when You choose the axes, so you can manufacture whichever gap you want. Unbounded by survey data it is wishful thinking — and the empty space is usually empty because nobody wants what is there.
A plot of a category's competing factors against the level each player offers on them, used to force raise, reduce, eliminate and create decisions.
Breaks when Examples are selected after the fact and there is no evidence on the method's forward success rate. Uncontested space is frequently uncontested because there is no demand in it.
A set of asymmetric strategies for brands that are not the category leader and cannot win by matching the leader's resources.
Breaks when Backfires if you are the leader — challenger posture is the market leader's most expensive mistake. Note also that Avis “We try harder” is the pre-framework archetype Morgan draws on, not a case of his method.
Non-name brand elements — colours, characters, sounds, shapes — that trigger the brand in memory, scored on fame and uniqueness.
Breaks when A new brand has no assets to score — investment first, measurement much later. And changing an asset resets everything accumulated in it.
Six mutually incompatible theories of how advertising produces its effect, set out together because the industry uses all six while claiming only one.
Breaks when It will not tell you what to do. Useful at the start of a brief, useless at the end of one.
Six properties — simple, unexpected, concrete, credible, emotional, story — shared by ideas that survive being retold.
Breaks when An evaluation tool posing as a generation tool. Chase all six at once and you achieve none.
A paired diagram mapping a customer's jobs, pains and gains onto a product's features, pain relievers and gain creators.
Breaks when Filled in around a table without talking to a customer, it becomes a register of assumptions — which is how it is usually filled in.
A seven-beat narrative template that casts the customer as hero and the brand as guide, applied to marketing copy.
Breaks when Produces the same shape in every sector: clarity bought at the cost of differentiation. Note also that no independently verified outcome case exists — every circulating figure traces back to a certified StoryBrand guide's own marketing.
A one-page hierarchy of a single overarching message, the pillars supporting it, and the evidence beneath each pillar.
Breaks when An internal alignment device. Copied straight into public-facing copy it produces corporate boilerplate.
A character, format or structure reused across successive campaigns so that recognition accumulates instead of resetting each time.
Breaks when Nothing compounds for a brand that repositions or changes agency often. Without accumulation the device is just a mascot.
Six levers of persuasion — reciprocity, commitment, social proof, authority, liking, scarcity — presented as a general toolkit for changing behaviour.
Breaks when It breaks when laboratory effects are assumed to transfer at laboratory size. Several of the underlying findings shrink sharply or fail to reproduce in field replication. It breaks harder when the principles are applied to brand advertising, where the audience is not deciding anything — scarcity aimed at someone who was not going to buy either way is a message with no mechanism. Compliance research generalises badly to advertising, which is mostly not asking for compliance.
The model that buyers pass in sequence through attention, interest, desire and action, with the population narrowing at each step.
Breaks when It breaks as a description of behaviour, which is the thing it claims to be. No empirical work has established that buyers move through these stages in sequence; measurement finds interrupted, non-linear, re-entered paths, and finds that most category buyers are in no stage at all at any given time. Treating it as real produces the characteristic errors: chasing “interest” as if it converts downward on a schedule, and building stage-gated budgets for a population that does not queue.
The repeated finding that the advertisement itself accounts for the largest single share of the sales effect — larger than targeting, reach, or where it ran.
Breaks when Two limits. The decompositions are run on campaigns that were already bought and distributed competently, so the finding says creative dominates among ads people actually saw — it does not say good creative rescues a campaign with no reach. And “creative quality” in these studies is largely an outcome-defined residual: what is left after the measurable variables are accounted for. Without a pre-test that predicts it, the claim edges toward circularity. Add that the major decompositions come from measurement vendors with a commercial interest in the conclusion, and this stays a Frame rather than a Law.
A screening procedure that ranks nineteen acquisition channels, cheap-tests three, and concentrates effort on the one that works.
Breaks when The list froze in 2015, and the method needs test budget and traffic to run at all. Very early on you need one bet, not a survey.
A model of acquisition in which the output of each cycle becomes the input of the next, as opposed to a funnel refilled from outside.
Breaks when Not every business has a loop. Forcing one produces a north-star metric that measures the wrong thing.
A five-stage funnel — acquisition, activation, retention, revenue, referral — used as shared vocabulary for diagnosing where users are lost.
Breaks when Assumes a linear one-directional flow, which flattens multi-touch B2B and pushes retention to the end where it does the least good. McClure's own funnel percentages carry his explicit “not actuals” disclaimer — never cite them as benchmarks.
The constraint that revenue per user and purchase frequency determine which acquisition channels a business can afford at all.
Breaks when Early on you do not know real ARPU or retention, so the arithmetic runs on optimistic estimates. Balfour gives no dollar thresholds — any you see attached to this are someone's invention.
The prescription, following from buyer-base structure, to reach as many category buyers as possible as continuously as possible rather than concentrating on heavy buyers.
Breaks when Inverts where the buyer universe is genuinely small and enumerable — a 300-account B2B list is a case where targeting really does pay.
A four-cell grid of growth options formed by crossing existing and new products with existing and new markets.
Breaks when It labels rather than directs. Knowing which box you are in says nothing about what to do inside it.
Two opposed rules for scheduling the same budget: buy enough repetition to cross a threshold, or spread continuous light coverage so an ad is present close to the purchase.
Breaks when Both break when treated as constants. The three-exposure threshold was a summary of the evidence available in 1979 and has never held as a universal number; recency assumes buying is continuous and randomly timed, which fails in seasonal and considered-purchase categories where there is a window and it is knowable. The deeper problem is that the two rules are not reconcilable — a planner citing whichever one supports the plan already chosen is using neither.
Measuring the seconds a person actually looks at an advertisement, instead of the opportunities-to-see that a media buy nominally delivers.
Breaks when It breaks on how the number is made. Attention is measured by eye-tracking on comparatively small opt-in panels, then modelled across platforms whose formats, screen sizes and viewing contexts differ enormously — so the cross-platform comparison, which is the use everyone wants, is the least supported part. Thresholds for what counts as attention are vendor-defined and not standardised, and the strongest evidence links attention to short-term sales measures rather than to long-term brand effects. Treated as currency it becomes a target, and the format that maximises measured seconds is not automatically the format that sells.
A four-question survey technique that derives an acceptable price range from judgements of too expensive, expensive, cheap and suspiciously cheap.
Breaks when Measures stated intention, not behaviour; systematically inflates the upper bound; ignores competitor prices and context. It gives you a range, never a price.
A survey technique that asks purchase likelihood at a series of set prices and builds demand and revenue curves from the answers.
Breaks when You choose the price points, so the answer comes from inside your own range. Still stated preference, and it does not model competitor response.
A method that infers the hidden weight a buyer places on each attribute and on price by observing forced choices between whole product bundles.
Breaks when Expensive, sensitive to design error, needs real sample. Past a certain attribute count respondents fatigue and the data degrades.
The maximum a rational buyer should pay: the cost of their best alternative plus the quantified value of the difference you provide.
Breaks when Strong in B2B, meaningless for emotional or status goods. And pricing at the full calculated value leaves the customer no reason to move.
A tiered price structure with deliberate constraints — fences — that stop customers in one tier from buying at another tier's price.
Breaks when If the difference between tiers is not perceived, everyone takes the middle one and total revenue falls. A tier without a fence is just a discount list.
An account of choice in which outcomes are judged as gains or losses against a reference point, with losses weighted more heavily than equivalent gains.
Breaks when Laboratory effect sizes shrink markedly in the field, and popular derivatives like the decoy effect replicate poorly. Trust the core asymmetry, not the ornaments.
Adding a deliberately worse third option to a choice set so that buyers move toward the option you want them to take.
Breaks when It breaks under replication. Large-sample and field attempts have repeatedly failed to reproduce the effect outside the original narrow stimulus designs, and where it does appear the size is far below what pricing-page advice implies. It also assumes buyers evaluate the set as a set, which is not how most purchases happen. The practical risk is not a null result — it is a live pricing page carrying a real option that some customers will actually buy, at a price you set to be unattractive.
The proportion of each joining group still active, plotted against time since joining.
Breaks when You cannot read flattening before enough time has passed. A curve drawn from three months of data is drawn from hope.
A family of probability models that predict, from purchase history alone, how many purchases a customer will make next and whether they have quietly stopped buying.
Breaks when The models assume a customer's underlying purchase rate is stationary — that it does not change. So they break precisely when you intervene: sustained promotion, a price change, a category shift, a competitor entering. They are a forecast of what happens if nothing is done, which makes them a strong baseline and a poor evaluator of your own campaign. They also explain nothing: a customer with a 4% survival probability comes with no reason and no lever.
A budget split, roughly 60% to broad-reach brand building and 40% to short-term activation, associated with the strongest long-term business effect.
Breaks when The ratio moves by category — the B2B figure is about 46/54, drawn from fewer than 50 cases, and the authors themselves call it tentative and warn against following it precisely. The database skews to large, mature, well-funded brands; a product starting from zero is not in it.
The gap between a brand's share of category advertising voice and its share of market, which predicts the direction and rate of share change.
Breaks when Share-of-voice measurement gets less reliable as media fragments, and the model ignores creative quality entirely — excess voice for a bad ad buys nothing.
The modelled carry-over of advertising effect after exposure, combined with the falling return on each additional unit of spend.
Breaks when Decay half-life and saturation point are estimated from data. With thin data they are effectively invented, and the model then confirms its own assumption.
A regression on aggregate historical data that estimates each marketing input's contribution to sales.
Breaks when Correlational. Uncalibrated by experiment it moves budget systematically to the wrong channel — and whoever builds the model can largely choose the answer.
A randomised experiment that withholds spend from matched markets or audiences in order to measure what that spend actually caused.
Breaks when Needs sample, time and patience. The statistical power to detect a small effect is absent from most budgets, so “no effect” usually means “we could not measure it”.
The ratio of a customer's expected lifetime gross profit to the cost of acquiring them, with payback period as the time taken to recover that cost.
Breaks when LTV is a forecast, and extrapolating it from early cohorts is the most common error in the field. Note also that the 3:1 ratio and the twelve-month payback rule are conventions, not findings — they are routinely presented as research.
Repeated survey measurement of a brand's presence in memory — recall, consideration, asset attribution — reported as a time series.
Breaks when Question wording determines the answer; mix prompted and unprompted recall and the series becomes meaningless. Change methodology and the entire history is void.
Rules that assign credit for a conversion across the touchpoints that preceded it.
Breaks when Observational, not causal. Benchmarked against experiments it systematically inflates cheap and search channels. It cannot substitute for an incrementality test, and it constantly is.
A single-question metric scoring likelihood to recommend on a 0–10 scale and reporting the share of promoters minus the share of detractors.
Breaks when The growth-prediction claim has been repeatedly refuted. Keiningham et al. examined 21 firms and 15,500+ customer interviews across five industries and found no support for NPS as the best predictor of growth — correlations were inconsistent and mostly non-significant, and in two of the three US industries Reichheld showcased, ACSI outperformed it. Useful as an internal trend; not an external truth.
The principle that an observed statistical regularity collapses once it is used as a target for control.
Breaks when It does not break. What breaks is the measurement system that forgot it.
A brand's share of category-related search volume, used as a fast and free leading indicator of market share.
Breaks when It is a proxy, and proxies drift from what they proxy. Search volume rises with newsworthiness as well as with demand, so a crisis and a successful campaign look the same. In categories where buying does not involve search, or where the brand name is also an ordinary word, the signal is mostly noise. And because it is cheap to move — brand-name bidding, PR spikes — it degrades the moment it becomes a target, which is the standard Goodhart path.
An experiment design that sizes a test to maximise total profit across the test and the rollout, instead of sizing it to reach statistical significance.
Breaks when It optimises the profit of one decision, so it deliberately accepts a higher error rate than a scientific test would — it will sometimes roll out the worse arm, by design, because the expected cost of that is lower than the cost of the larger test. That makes it the wrong instrument when the result has to generalise, feed a model, be reused across markets, or be defended to someone outside the team. It also needs a prior on the effect size, and a badly chosen prior moves the recommended sample size a long way.
Nothing matches that. .