Why last-click attribution quietly killed your brand budget
18 November 2021
Marketing mix modelling needs aggregate spend and outcome data over a reasonable history, not a perfectly governed data warehouse. Gaps, inconsistent naming, missing weeks and changed definitions are all workable. They are modelling problems with known treatments, not blockers.
The belief that they are blockers has a cause. Platform-based providers need clean, structured inputs because their pipelines are automated. When the answer to a data quality problem has to come from software, the software's requirements become the project's requirements.
| Requirement | Why | What we can work around |
|---|---|---|
| Two to three years of history | You need enough variation in spend to estimate a response curve, and enough cycles to separate seasonality from effect | Shorter histories can work at higher granularity, with wider error bars stated honestly |
| Weekly granularity | Monthly data hides the lag structure that makes the model useful | Monthly can be modelled, but the diminishing returns estimates get soft |
| Spend by channel | The independent variables | Inconsistent channel taxonomy over time is normal and mappable |
| An outcome series | Sales, joins, leads, retention, whatever the business is actually managed on | Multiple outcomes are better than one; churn is frequently the more valuable model |
Note what is absent from that list: a customer data platform, resolved identity, a tag audit, or a completed data governance programme. Those are worth having for other reasons. They are not prerequisites here.
A quarter of missing spend data is a gap to be handled explicitly, not a reason to abandon the exercise. What matters is that the treatment is documented and its effect on confidence is reported rather than smoothed over.
Almost every organisation has restructured its reporting at some point. Reconciling two taxonomies is manual work of a few days, and it usually surfaces useful history in the process, including campaigns nobody currently on the team remembers running.
Ideally you have a brand health series. If you don't, Share of Search is a usable proxy for brand demand: it is external, free, updates weekly, and research popularised by Les Binet found it tracks and can lead market share.
Worth naming, because it is the case that most often causes real trouble. Data that looks immaculate has usually been through a reporting layer that already applied attribution rules, deduplication or modelled conversions. Feeding that into a model means modelling somebody else's assumptions. We would rather have the messy source.
Because we are not a platform, we can account for issues inside the data rather than requiring them to be fixed first. In practice that is also why the analysis usually extends well past the original scope. The interrogation needed to handle the mess turns up questions worth answering.
Why we work the way we doA data readiness programme is a multi-year commitment with no interim answer to the question your CFO is asking now. Meanwhile the decisions continue to be made: on last year's split, on the agency's recommendation, or on judgement.
Our preference is to model with what exists, deliver a working answer in weeks, and let the modelling itself tell you which data investments are worth making. The gaps that materially widen your error bars are the ones to fix. The rest can wait, and quite often should.
Tell us the decision you are trying to make. If we are not the right people, we will say so and point you somewhere better.
Start a conversation