Big Data and Public Policy: What It Changed
The question is not whether administrative and behavioural data is useful. It is which decisions are actually constrained by a lack of measurement, because most are not.
A decision is improved by better data only if measurement was the binding constraint. Many public decisions are constrained by budget, by legal authority, by political feasibility or by disagreement about objectives. Adding a dashboard to those changes nothing. The first question for any data-for-policy project should be which specific decision changes if the number comes back different, and a surprising proportion of projects cannot answer it.
Where it has worked
Timing. Official statistics are accurate and slow. Digital sources are noisier and immediate. Where a response must happen before the accurate number exists - mobility after a disruption, demand shifts during a shock - the noisy fast estimate is genuinely better, because the alternative is not a better estimate but no estimate.
Coverage of the hard-to-survey. Populations that surveys systematically miss are sometimes visible in administrative or transactional data. This works when the coverage gap in the new source does not coincide with the gap in the old one, and it fails when both miss the same people.
Geographic and temporal resolution. Aggregate statistics conceal variation that matters for targeted intervention. Data at the level of a district and a week supports decisions that a national quarterly figure cannot. This is the most reliable class of gain.
Where it fails
The instrument moves. A platform changes its ranking, an app changes its default, a process changes its recording rules, and your series develops a break that looks like a change in the world. Digital sources are not designed for measurement, so nobody tells you when the instrument changed.
The measure becomes the target. Once a proxy is used for allocation, behaviour adjusts to the proxy and it stops tracking what it proxied for. This is not a subtle risk; it is the default outcome and should be assumed unless there is a reason it will not happen.
Selection is mistaken for signal. The best-documented case remains search-based influenza estimation, which worked well initially and then drifted badly. The failure had two components that recur: the search behaviour was partly a response to media coverage rather than to illness, and the underlying platform kept changing what search suggestions it offered. The lesson was not that the approach was worthless - combined with conventional surveillance it improved on either alone - but that a proxy validated once stays validated only as long as the process generating it holds still.
The evaluation problem
Large observational datasets make correlations abundant and causal inference no easier. The scale that makes big data attractive also makes every correlation statistically significant, so significance stops carrying information and effect size and identification carry all of it. Policy questions are causal questions, and answering them still requires a design: randomisation where possible, and a credible identification strategy where not. The methodological standards are the ones set out in validation and calibration.
Simulation as the complement
Data describes what happened under the policies that were in place. Policy questions concern what would happen under policies that were not. That gap cannot be closed by more observation, only by a model of the mechanism, which is the practical argument for agent-based approaches in policy analysis: not better forecasts, but the ability to ask counterfactual questions at all.
Distributional blind spots
Data infrastructure is unevenly distributed, and the populations best measured tend to be the ones best served already. Building allocation rules on the available data therefore embeds existing coverage gaps into the allocation, which is a mechanism for reproducing inequality that operates entirely through apparently neutral technical choices.