Notes on:
Measurement without Random Validation Samples: Multiple Measurements and Data Fusion
5 August 2026
econometrics · measurement · data fusion
Talk · Paper · Slides · Transcript
Written by Fable 5
Part of NBER Summer Institute 2026 Methods Lecture: Estimation and Inference with AI-Generated Data
Ashesh Rambachan (MIT), part 4 of 4 of the NBER Summer Institute 2026 Methods Lecture “Estimation and Inference with AI-Generated Data” (with Melissa Dell), July 30, 2026. Slides: ash2.pdf; the companion background reading: Ludwig, Mullainathan & Rambachan’s “Large Language Models: An Applied Econometric Framework”. Timestamps refer to the lecture video.
Dell’s framework needs two things: a measurement process you’d actually defend, and a random validation sample drawn from your population. Rambachan spends the closing lecture on what happens when you can’t have either — and the answer, in both cases, is that you end up making assumptions about the behavior of frontier AI models, which is an empirical question you can go and check. He checks. The assumptions mostly fail.
Case one: there is no existing measurement to scale. The pushback he hears most is from text people: in 2026, why privilege trained human annotators once the rubric is written? GPT-3.5 aced the GRE in 2023; a model won IMO gold last summer. Grant the point, he says, and hand the rubric to a frontier model instead. You immediately discover the rubric doesn’t pin down the measurement, because you still have to pick a model family and a prompting strategy — plus the sampling randomness and the silent model updates that economics journals’ data editors flagged just last week. Does that matter? He takes ~10,000 congressional bills whose policy areas political scientists hand-coded through a heroic semester-long annotation effort, relabels them with various GPT models and prompts, and runs the obvious regressions:

The same exercise on 140 years of congressional immigration speeches (labeling tone as pro-, neutral, or anti-immigration) gives the same result: the estimated partisan tone gap swings in magnitude and sign across models and prompts. And revisiting the 2024 exercise with GPT-5 a year later didn’t make the problem go away.
One escape is to declare that the model’s output is the construct — the measurement process is “this model under this prompt.” Rambachan finds that road unattractive for a reason that is really a question about the sociology of the field: are researchers using Claude and Gemini estimating different estimands? Do we revisit every empirical paper when a new model ships? (He notes the promising work on building better harnesses — the scaffolding around a model — to reduce prompt sensitivity, and brackets it.)
The other escape is classical: treat the construct as inherently latent and each model-prompt combination as one of many repeated measurements of it. Every econometrician knows what comes next — with measurements that are conditionally independent given the latent construct, plus a mild informativeness condition (true positive rate exceeds false positive rate), the counting works out (you learn probabilities, you need unknowns) and the latent base rate is identified. This is the workhorse of the measurement-error literature. The catch is that when the measurements are LLM outputs, conditional independence is an assumption about frontier AI behavior — it says errors are unrelated across models and prompts once you condition on the truth. Frontier labs share architectures, pre-training corpora, and RLHF recipes. So: go look.
The computer scientists already did, under the banner of algorithmic monoculture: across many public models and benchmarks, the rate at which two models give the same wrong answer far exceeds chance — and, disquietingly, correlated errors are worse on better-performing models. Rambachan replicates the exercise on economics tasks, testing pairwise whether errors are independent conditional on the hand-coded ground truth:

The verdict is blunt: identification results resting on conditional independence “should not be naively applied to LLM-generated measurements.” But he refuses the nihilistic reading — errors are correlated, yet not arbitrarily correlated, and the structure differs across model families versus prompts within a family, which is exactly the kind of thing a partial-identification approach can exploit. His constructive claim is that this whole exercise is a template: economics has a huge literature deriving identification from measurement-error assumptions; each such assumption is now a testable claim about model behavior, and benchmarking is how you discipline it. Except the benchmarks that exist are the wrong ones — built for math and coding, scored on overall accuracy, which is a misleading guide to downstream estimation bias. Building measurement benchmarks for the social sciences is, he argues, essential and undone.
Between the two halves comes the lecture’s best paragraph, a defense of Dell’s framework against a misreading. The validation-sample approach is not about scaling the compromise you settled for before AI; it’s about scaling the measurement you would actually defend — a documented, adjudicated rubric applied by the researcher, not by “the undergrad RA that you were able to cheaply hire.” Because your best measurement no longer has to cover the corpus, you can afford to make it genuinely good on a small sample and let AI carry it across the rest. The scarce input is now high-quality measurement itself. And the historical parallel gets restated with force: when running regressions became cheap, econometrics stopped worrying about inverting and started worrying about credible identification. Same move here — invest in the measurement, then let it rip.
Case two: the measurement exists, but in another sample. This is the remote-sensing and digital-trace world: you want household consumption or forest cover, the survey or census that measures it properly was run in another region or another year, and all you have in your sample is a satellite-based prediction. First, does prediction quality travel? Proctor and coauthors benchmarked ~115 socioeconomic and environmental variables predicted from the public MOSAIKS satellite embeddings, and performance degrades sharply with geographic distance from the training sample across every outcome class. Rambachan replicates it in about a week with Codex agents — living standards (a 14-component DHS index) trained on Kenyan clusters and transferred to other parts of Kenya, to Ethiopia, and across time; forest cover in Uganda, the setting of Jayachandran and coauthors’ early deforestation experiment — and finds substantial transfer error in space and time, for both MOSAIKS and the newer AlphaEarth embeddings. So you cannot just assume the predictor behaves the same in the two samples; you have to assume something specific.
Two natural candidates. Outcome stability: the distribution of the true outcome given the prediction and covariates is the same across samples — i.e., the predictor has the same calibration curve in both. Believe it, and you’re essentially in a surrogacy framework with known identification results. Measurement stability: the distribution of the prediction given the true outcome and covariates is the same — i.e., the same true-positive and false-positive rates in both. (Intuition: crop burning produces the same visual signature in satellite imagery in one part of Malawi as another.) Believe that, and Rambachan’s work with Rahul Singh and Davide Viviano identifies the target through a conditional-moment/IV-style problem.
Then the elegant sting. Run Bayes’ rule from one to the other and you find each implies the other only under a knife-edge cancellation involving the base rates in the two samples. Formally: both can hold only if the base rates agree or the predictor is error-free — and in a data-combination problem you chose the auxiliary sample precisely because it’s elsewhere or earlier, so base rates differ, and the benchmarks just told you predictors are imperfect. Readers who know the algorithmic fairness literature will have recognized both assumptions as canonical notions of a predictor “behaving the same” across subpopulations, and the incompatibility as a restatement of the decade-old fairness impossibility results. Economics has rediscovered, in the guise of data fusion, a theorem about calibration and error-rate balance that computer science proved while arguing about recidivism scores.
So you must pick an assumption and defend it — and the stakes cut both ways. A semi-synthetic exercise on deforestation data from the DRC, Brazil and Indonesia, constructed so that measurement stability holds, shows estimators exploiting it are approximately unbiased across data-generating processes while estimators that don’t can be wildly off; get it right and you also tighten standard errors by 20–40% relative to alternatives. His summary of the econometrics-police position: “If you choose the wrong assumption, you’re going to get wrong answers. So choose the right assumption, think hard.”
He closes past six o’clock with the frame he uses to end his own course. The crowning achievement of the last twenty-five years of econometrics was the causal inference revolution, which gave the field new ways to analyze data and thereby new questions it could ask. The claim of these four lectures is that the same thing is now being built for measurement: hypothesis generation to widen the aperture on what we even think to measure, and validation-sample debiasing to scale the measurements we’d defend. “Measuring better allows us to see the world in new ways.” Then, having spent forty minutes demonstrating that frontier models make correlated errors, transfer badly, and quietly change under you, he offers his inbox and tells the room to go out and measure interesting things.