9.7
Test & Roll
An experiment design that sizes a test to maximise total profit across the test and the rollout, instead of sizing it to reach statistical significance.
Elea McDonnell Feit and Ron Berman, Marketing Science (2019)
What it does
Reframes the sample-size question from “how many do I need to be sure?” to “how many can I afford to show the worse option to?”. Conventional power calculations routinely prescribe tests larger than is profitable, because every unit in the losing arm is real money spent on the treatment you are about to abandon. Test & Roll gives a sample size and a decision rule that account for that cost, and its answer is often to test smaller and simply pick the winner.
When it breaks
It optimises the profit of one decision, so it deliberately accepts a higher error rate than a scientific test would — it will sometimes roll out the worse arm, by design, because the expected cost of that is lower than the cost of the larger test. That makes it the wrong instrument when the result has to generalise, feed a model, be reused across markets, or be defended to someone outside the team. It also needs a prior on the effect size, and a badly chosen prior moves the recommended sample size a long way.
Case
Feit and Berman derive the profit-maximising sample size for a two-arm test and show that it is typically far smaller than the size a conventional hypothesis test would require, with the difference being profit left on the table by over-testing. Replication code for the paper is published openly.
Feit & Berman — Test & Roll: Profit-Maximizing A/B Tests, Marketing Science ↗Diagram — not yet drawn
Profit plotted against test size as a curve that rises then falls, with the significance-driven sample size marked well past the peak and the shaded area between the two labelled as the cost of certainty.
In the wild
Unvetted · not part of the tier assessment
What has been written about this tool in the last twelve months. Machine-retrieved and unchecked — everything above this line was checked.
Loading…