Measurement and Identifiability
Before asking whether a go-to-market system works, ask whether 'works' can be known from the data at all.
Before asking whether a go-to-market system works, ask whether “works” can be known from the data at all.
A company I know had a channel it was certain about. The branded search ads, the ones that showed up when someone typed the company’s own name into a search engine, returned customers at a cost so low that the team treated the spend as close to free money, and the attribution data agreed, crediting those ads with a steady stream of conversions at an enviable return. Then, half by accident, the ads were switched off for a stretch, and the conversions barely moved. The people who had been clicking the branded ads, it turned out, were people already looking for the company by name, people who would have found their way to the site through the ordinary search result sitting just below the ad, and the ads had been collecting credit for arrivals that were going to happen regardless. The channel that looked like the best-performing thing the company did was, in the part that mattered, buying customers it already had.
What made the discovery so unsettling was that nobody had been negligent. The team had watched the right dashboard, trusted accurate data, and drawn the obvious conclusion from it, and the conclusion was still wrong, because the dashboard answered a question subtly different from the one that mattered. It told them which conversions followed the ads. It could not tell them which conversions the ads caused, and the difference between those two questions, invisible on any dashboard, was most of the channel’s apparent value.
This is one of the most important and least understood facts about measurement in go-to-market, and it is not a story about bad tracking. The tracking was fine. The numbers were accurate. Every conversion the attribution system credited to those ads had really happened, and really followed a click on the ad. The problem was deeper than accuracy, and it sits underneath nearly every measurement the field makes, which is that the question the company actually cared about, how much demand the ads caused, is a question about causation, and causation cannot be read off observational data no matter how clean the data is. A field that wants to know whether its systems work has to confront this, because the thing it most wants to measure is the thing its data, by itself, cannot tell it.
The question is causal, and the counterfactual is missing
When you ask whether a channel works, what you mean, whether or not you say it this way, is how much demand it caused, how much more demand exists because the channel ran than would have existed if it had not. That comparison, between what happened and what would have happened otherwise, is the definition of a causal effect, and it has a structure worth making explicit, because the structure is where the trouble lives.
For any unit, a person, a market, a period, there are two potential outcomes. There is the outcome if the action is taken, which I will write , the demand you get if the channel runs. And there is the outcome if the action is not taken, , the demand you would have gotten if it had not. The causal effect of the action for that unit is the difference between them, , and what you care about across the whole population is the average of that difference, the incremental effect,
This is the real quantity, the thing the word works is reaching for, the demand the action actually caused. And here is the difficulty that no amount of data removes. For any given unit you can only ever observe one of the two outcomes. If the channel ran, you see and never the that would have happened in the world where it did not run, because that world did not occur. The other outcome is a counterfactual, and counterfactuals are never observed, which means the quantity you most want is built out of something you can never directly see. This is sometimes called the fundamental problem of causal inference, and it is fundamental in the strict sense, a permanent feature of the situation rather than a limitation of current tools, and it is the reason measurement in go-to-market is so much harder than it looks.
It is worth sitting with how strange and absolute this is. The missing outcome is missing as a matter of logic, not of effort. You cannot run the world twice, once with the channel and once without, and compare, because the world only happens once, and whichever version you ran is the only one you will ever observe. Every claim about what a channel caused is therefore a claim about a world that did not happen, the world where you held the channel back, and that world has to be reconstructed somehow, because it can never be witnessed. All of causal inference is, at bottom, the science of reconstructing the world that did not happen, and the quality of any causal claim is the quality of that reconstruction.
What attribution actually measures
Set the counterfactual against what attribution systems actually do, and the gap becomes stark. An attribution system watches what happened, sees that a conversion followed a click on a channel, and credits the channel with the conversion. Last-touch attribution credits the last click, other schemes spread the credit across the touches, but all of them work from the same raw material, the observed path of people who converted, and all of them are therefore measuring association, the demand that co-occurred with the channel, rather than , the demand the channel caused.
The branded-search story shows exactly how these come apart. The attribution system saw conversions following clicks on the branded ads and credited the ads, which is correct as a statement of association and wrong as a statement of cause, because most of those people had a high , a strong chance of converting even with the ads switched off, since they were already searching for the company by name. The incremental effect of the ads was small, because and were nearly the same for the people the ads reached. The attributed credit was large, because all the conversions co-occurred with the ad. The number the company steered by was the attributed credit, and it bore almost no relation to the incremental effect, which is the only thing that should have mattered.
What makes this gap dangerous is that it runs in a consistent direction depending on the kind of channel. Channels that intercept people who are already on their way to converting, branded search, retargeting that follows people who already visited, much of the bottom of the funnel, systematically show attributed credit far above their incremental effect, because they touch people with high and harvest demand that already existed. Channels that create demand earlier, that reach people long before they convert and influence them in ways no click captures, systematically show attributed credit far below their incremental effect, because the demand they caused converts later through some other channel that collects the credit. The field’s standard measurement therefore over-credits the harvesting channels and under-credits the creating ones, in a predictable direction, which means a field that allocates by attribution will systematically overfund the channels that capture demand and starve the channels that create it, while believing it is following the data.
Retargeting is the purest example of the trap, the ads that follow someone around after they have visited a site. By construction they reach people who have already shown enough interest to visit, a group with a high chance of returning on their own, so the attributed credit is enormous and the incremental effect is often small, because many of the people clicking the retargeting ad were coming back regardless. A company watching attribution sees retargeting as one of its best channels and pours money into it, and the money is largely buying conversions it already had, the branded-search mistake wearing slightly different clothes. The pattern repeats anywhere a channel selects for people already close to converting, which is precisely where attribution looks most impressive and incremental effect hides lowest.
A toy example fixes the scale of the problem. Suppose attribution credits a retargeting channel with a thousand conversions in a month, and the team, seeing that number, counts the channel a triumph. Then a holdout runs, half the eligible users randomly held back from the ads, and the comparison shows the exposed group converting only a little more than the held-back group, enough to imply a true incremental effect of around a hundred and fifty conversions, with the other eight hundred and fifty arriving regardless. The attributed credit was off by nearly sevenfold, and every dollar spent chasing the attributed thousand was being justified by eight hundred and fifty conversions the channel never caused. The numbers here are invented, but the size of the ratio is not unusual in the published holdouts, which is what makes the gap a budget problem rather than a rounding error.
Identifiability, the question beneath measurement
The deep version of all this is a question that the field has never learned to ask, the question of identifiability. To identify a quantity is to be able to recover it, even in principle, from the data you have, and the sharp and uncomfortable fact is that some quantities are not identifiable from observational data at all, no matter how much of it you collect. Identifiability is a property of the question and the data together, prior to any estimation, and a quantity that is not identified cannot be estimated, because there is nothing in the data that pins it down.
The incremental effect is the case that matters. The naive way to estimate it is to compare the outcome among the treated to the outcome among the untreated, the people the channel reached against the people it did not,
This recovers the true only when the treated and untreated groups would have behaved the same in the absence of treatment, when the untreated group is a valid stand-in for what the treated group’s would have been. In observational data that condition almost never holds, because the people a channel reaches differ systematically from the people it does not, and they differ precisely on their propensity to convert. The branded-search ads reached people already searching for the brand, a group with a far higher than the general population, so comparing them to anyone else overstates the effect. The bias is built into the structure of who gets treated rather than being noise that averages out, and collecting ten times as much data only buys you a more precise estimate of the wrong number.
This is the point that changes how you see the whole problem. The failure of attribution is not a measurement error to be reduced with better tracking or more data, it is a failure of identification, a case where the quantity is not recoverable from the data at hand even in principle. No dashboard, no matter how sophisticated, can extract a number the data does not contain, and the incremental effect is, in observational data with selection, exactly such a number. The field has spent enormous effort on the tracking and almost none on the identification, which is the part that actually determines whether the number can mean what it claims.
This is the hardest idea to absorb, because everything in the culture of data pushes the other way, toward the belief that more data and better tools eventually reveal the truth. For questions of association, that belief is sound, and more data does sharpen the estimate. For questions of causation under selection, it is false in a specific and permanent way, because the bias lives in the gap between who was treated and who was not, rather than in the sample size, and that gap is still there in a sample a thousand times larger. A biased estimator converges to the biased answer as data grows, more and more confidently, and never to the truth, which is the worst of both worlds, a wrong number wearing the credibility of a large sample.
How the effect becomes knowable
Identifiability also tells you the way out, because it specifies exactly what would make the incremental effect recoverable, and there are only a few routes.
The first and cleanest is experiment. If you randomly assign who is exposed to a channel and who is held back, randomization makes the treated and untreated groups statistically the same in everything, including their unobserved propensity to convert, so the held-back group’s outcome is a valid estimate of the treated group’s , and the difference between the groups identifies . This is what a holdout test, a geographic experiment, a randomized incrementality study actually does, and it is why these methods are the gold standard. They manufacture the missing counterfactual by force, creating a group whose you get to observe, and they are the only general way to know an incremental effect with confidence. The branded-search company learned the truth the moment it ran the experiment of switching the ads off, which was a crude holdout, and the crude holdout told it more than years of accurate attribution had.
The geographic experiment is the workhorse version of this for channels that cannot be randomized at the level of the individual person. You run a channel in some regions and hold it back in others, chosen to be comparable, and the difference in demand between the on regions and the off regions estimates the incremental effect, with the off regions serving as the observable counterfactual. It is coarser than a person-level randomization, and it costs real money, because the held-back regions are deliberately under-served for the length of the test, and that cost is exactly why the field avoids it and exactly what it buys, a number that means what it claims. The expense of the holdout is the price of knowing, and a field that will not pay it has chosen, in effect, to keep guessing in a way that feels like measuring.
The second route is assumption. In some situations you can identify a causal effect from observational data by assuming that you have accounted for everything that makes the treated and untreated groups differ, that there is no unobserved variable driving both the treatment and the outcome. This is a real method, and sometimes it is the best available, and the crucial discipline is that the assumption has to be stated, defended, and treated as the load-bearing thing it is, because the estimate is only as good as the assumption, and the assumption of no unobserved confounding is usually fragile and often false in exactly the cases the field cares about. An effect identified by assumption is a conditional truth, true if the assumption holds, and honesty requires carrying the condition along with the number.
The third route is to accept that the effect is not knowable, which is a real and underused answer. Some quantities, with the data and the experiments available, cannot be identified, and the disciplined response is to say so, to mark the number as unknown rather than to report a biased estimate as though it were the effect. A field that could distinguish what it knows by experiment, what it believes by assumption, and what it cannot know at all would be far ahead of one that reports every number with the same false confidence, because knowing the epistemic status of a number is part of knowing the number.
The hypothesis
Let me state the claim as a hypothesis sharp enough to test and to be wrong.
The hypothesis has two parts. The first is that standard attribution, last-touch and its observational relatives, is a biased estimator of the incremental effect except under conditions, randomization or genuine absence of confounding, that rarely hold in practice, so that the numbers the field steers by are measures of association standing in for a causal quantity they do not recover. The second is that the bias is directional and predictable, overstating the incremental effect of demand-harvesting channels that select for high-propensity users and understating the incremental effect of demand-creating channels whose influence converts elsewhere, so that allocation by attribution systematically overfunds harvesting and starves creation.
These are testable, and the test is the experiment itself. Run real holdouts on a range of channels, measure the incremental effect directly, and compare it to what attribution credited. The hypothesis predicts large gaps, and predicts their direction, with the harvesting channels showing measured incremental effects well below their attributed credit and the creating channels showing the reverse. The branded-search company’s accidental experiment is one data point in exactly the predicted direction, and the published cases that exist, where firms have run disciplined holdouts against their attribution, point the same way, often dramatically. If you ran these experiments and found attribution closely matching incremental effect across channels, the hypothesis would be wrong, and the field’s measurement could be trusted as it stands, which would be a relief and a surprise.
What it changes
Everything in the engineering of go-to-market that depends on knowing whether something works depends on this, which is to say nearly everything. The control loops that steer a system, the loop gains that decide whether growth compounds, the channel curves that govern allocation, all of them are computed from measurements of effect, and if those measurements are biased estimators of association standing in for cause, then the whole apparatus is steering by a number that does not mean what it claims, confidently and in a consistent wrong direction.
There is a particular trap in this for the data-driven team, the one that prides itself on deciding by the numbers, because that team is the most exposed of all. A team running on intuition at least knows it is guessing and stays alert for being wrong. A team running on biased measurement believes it is being rigorous, and trusts the number precisely because it came from data, and so steers harder and more confidently in the wrong direction than the intuitive team ever would. The appearance of rigor, without the identification that rigor requires, is more dangerous than open guessing, because it removes the very doubt that would otherwise keep a team watching for the failure. Measurement determines whether the real work can be known to be working, which makes it foundational rather than a reporting layer sitting on top, and a discipline that cannot identify its own effects is not yet able to know itself.
A field that took this seriously would reorganize around the experiment, treating the holdout and the geographic test and the randomized incrementality study as the central instruments they are, the only general way to manufacture the counterfactual that causation requires. It would treat attribution as a useful description of association while refusing to mistake it for a measure of cause. It would state its assumptions whenever it inferred an effect without an experiment, and it would mark the truly unknowable as unknown. And it would understand that this is the foundation rather than a technical nicety, because a discipline is, among other things, a way of knowing whether what you did had the effect you think it did, and a field that cannot answer that question has not yet earned the name.
There is a final turn here that reaches past technique. Once you accept that many of the effects you act on cannot be fully verified, that you are steering powerful systems on numbers that are estimates under assumptions, a question arrives that measurement alone cannot answer, which is what you owe to the people on the other end of a system you are operating partly in the dark. The systems go-to-market builds act on real people, and they act at scale, and the engineer who builds them is making choices whose full effects can be neither seen in advance nor cleanly measured after, which is the condition under which questions of responsibility stop being abstract and become the heart of the matter.
References
- On potential outcomes and the fundamental problem of causal inference, see Paul W. Holland, “Statistics and Causal Inference,” Journal of the American Statistical Association (1986), and the broader Rubin causal model; for the structural account, Judea Pearl, Causality (Cambridge University Press, 2009).
- On identifiability as a property of question-and-data prior to estimation, see standard treatments in causal inference and econometrics.
- On the gap between attributed and incremental effect in paid search, see Thomas Blake, Chris Nosko, and Steven Tadelis, “Consumer Heterogeneity and Paid Search Effectiveness: A Large-Scale Field Experiment,” Econometrica 83, no. 1 (2015): 155–174, and the broader incrementality-testing literature.