Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # Measurement without Random Validation Samples: Multiple Measurements and Data Fusion, Ashesh Rambachan, Massachusetts Institute of Technology Authors: Discussant: None Video: https://www.youtube.com/watch?v=ydYXe2FxErk&t=8132s ## Talk (02:15:32 – 02:54:22) [02:15:38] I'm getting sleepy myself. Um but you know, there's a ton of material here. This would be stuff that we would probably cover over several months in a PhD course. So, slides are online, tons of references, lots of resources. Please [02:15:51] reach out if you have questions. Um so, you know, the last uh 40 minutes or so I want to talk about, you know, what happens if we fall our find ourselves outside of the case um that Melissa um was just dis- discussing at length. Um and so let me just sort of talk a little [02:16:05] bit briefly about how I think about this validation sample framework um at a very high level. So, you know, in a lot of settings, you know, we as researchers we're going to collect some unstructured data, documents, images, videos, and link them to some covariates. And the reason we're interested in this [02:16:19] unstructured data is we think that it expresses some construct that we're interested in studying. Um and that construct might be defined by an existing measurement process that in my slides I'm going to denote by F star. [02:16:31] And then we would like to report some downstream parameters that are defined over the joint distribution of VI star and WI. This measurement process is something that the researcher has to choose themselves. Uh and I think of it is what the researcher would use if they [02:16:44] were truly unconstrained in terms of time and resources. So, in text analysis, often times that looks like, well, we would like to write a very careful and extensive rubric that defines the construct and adjudicates edge cases. And then in ideal scenario, [02:16:58] we would have experts or a team of experts label each document in light of this rubric often in short sessions to prevent fatigue from introducing error. [02:17:07] Or in environmental/development economics, perhaps the ground truth, the best-case measurement process is that we would send surveyors into remote forests at high frequency to measure tree cover. [02:17:17] We would send surveyors out to remote villages in order to ask people extensive consumption surveys. Or we would blanket physical fields with uh you know, actual monitors in order to measure pollution. [02:17:29] Now, the core problem we face in doing this sort of research is there's a measurement bottleneck here. Applying our existing measurement process across the full sample is just infeasible. We could not have a team of experts get a bunch of macroeconomists in the room to carefully read every single newspaper [02:17:44] article that has been published over several decades. For cost reasons, we can't send surveyors into remote forests at very high frequencies or blanket fields in our experiment with physical monitors. [02:17:54] So, this existing framework on using validation data asks the question, well, how could we use machine learning or AI-generated measurements that I'm going to denote by RI to solve this measurement bottleneck. [02:18:05] And as Melissa you know really nicely explained the solution so far is the following. Let's collect our machine learning or AI generated measurement on our full sample of interest. On a random validation sample let's collect the existing measurement we would like to [02:18:20] study and let's use that validation sample to debias the machine learning or AI generated measurements. And this is really a beautiful idea. It's you know built on a lot of old ideas in measurement error and survey sampling and it's been really recently re [02:18:34] resuscitated in a lot of work at the intersection of machine learning and AI and Melissa's paper with Jacob is a great introduction to the ideas in this literature. [02:18:42] Now why is this solution so attractive? The reason the solution is very attractive is that frontier machine learning and AI models are extraordinarily complicated. If you think about a language model that's billions and billions of parameters trained on data that you have never seen [02:18:56] nor will ever touch and those models can make errors. So we would like to correct for measurement error in their outputs without having to make strong assumptions about properties of these outputs and this is a framework that allows us to do that. [02:19:10] So the payoff of having validation samples is that we can recover the estimate that we're interested in which is defined over the joint distribution of V star and WI while only paying measurement costs on a subset of our [02:19:23] sample on the random validation sample. Of course to execute this framework there are two requirements here. [02:19:29] The first is the researcher must choose an existing measurement that they would like to scale over their sample. They have to define this object F star. [02:19:38] And the second is that the validation sample must be collected from the researcher's population of interest. [02:19:44] So I want to talk about in this lecture is you know what happens when these two requirements fail. [02:19:48] In the first case what happens if there is no existing measurement process that we would like to scale? I think in you know going over the last couple of years in presenting these ideas, uh people push back and make this argument most often in text applications where they would [clears throat] like to label [02:20:02] documents with large language models. So, I'm going to illustrate this idea in that setting. And the second is going to be one in which there is in fact an existing measurement that we care about and would like to defend, but it exists in some other sample. Um and this most [02:20:15] often arises in applications involving remote sensing with satellite imagery um or digital trace data where there is existing measurements from surveys that we may have run or existing census censuses that cover another region or period than the one that we're studying [02:20:29] at hand. So, in both of these points, there going to be some key ideas I want you to take away and you know, there's going to be some details. Don't worry about if they go too fast. We'll slides will be available, but the first one is if there is no existing [02:20:42] measurement uh that we'd like to scale, we're going to have to fall back and start introducing assumptions about measurement error. And what I want you to keep in mind is that when we start making assumptions about measurement error in these settings, we're really making assumptions about the behavior of frontier AI models. So, I'm going to [02:20:57] illustrate that in the context of one example where the construct that we're interested in studying is latent, but we might want to treat ML or AI as offering us many repeated measurements of that latent construct. And then I want to revisit some classical identification results from the measurement error [02:21:12] literature using what we know about frontier model behavior in computer science and AI. [02:21:18] In the second case, when there is an existing measurement we care about, but it exists in another sample, we're going to face a data combination problem. And what I want you to take away is that identification is going to turn on what you're willing to believe is stable [02:21:32] between your sample and the auxiliary sample where you have the existing measurement and the AI output. And what I want to point out is that two very natural assumptions we could write down here both work quite well if they are true are actually going to be incompatible with each other outside of [02:21:45] knife edge cases. So, this is going to be a place where you'll really have to pick an assumption and try try defend it. And we'll talk about how you might approach that problem. [02:21:54] So, let's think about this first case where we worry that there may not be an existing measurement we want. So, again, we want to think about this in the context of text where the measurement process, the best-case scenario, typically proceeds in two steps. Uh step one, as Melissa was articulating, is one [02:22:08] of definition. We, the research team, would want to draft a very detailed rubric for the construct that tries to adjudicate edge cases and um surface surface edge cases, possible disagreements, and iterate until we're satisfied we have, you know, worked out the kinks. [02:22:21] Then in step two, we'll take that rubric and train annotators, ideally experts in the domain, to apply that rubric across the corpus. So, what the solution Melissa presented so far I would say is to incorporate machine learning and AI into this process. Let's invest in step [02:22:35] one, create a detailed rubric that we would care about, then only run step two on a random subsample. If we do this appropriately, we can still recover the S-demand defined by V-star. [02:22:46] Now, I think a question definitely in my mind, maybe lurking in the back of your mind, is uh in 2026, why should we privilege trained human annotators once the rubric is defined? [02:22:57] Uh after all, you know, in 2023, GPT-3.5 was already able to ace the GRE. Last summer, it was getting a gold medal in the IMO. Uh and it seems like every single month there is a new model. And if you're on Twitter, you're seeing [02:23:10] impressive benchmarks about them. So, if Frontier AI can accomplish these feats, why are we trying to privilege human annotators in this process once we have defined the rubric? [02:23:19] So, I want to now grant that point for a second. Uh and let's just imagine we took the rubric we wrote down and just presented it to Frontier models rather than train annotators. [02:23:28] And what we, you know, it's important to keep in mind is that as Melissa was articulating, the rubric alone does not pin down the measurement that's going to be produced, say by a Frontier language model. There are many choices that the researcher still has to make. First, [02:23:41] what model do I use? Do I go to the Open AI family? Do I go to the Anthropic family? Do I go to Gemini? Do I go to G DeepSeek or some open source model? [02:23:49] And then once even when I've committed to a model, there's a question of what's the prompt engineering strategy going to look like? What's going to go into the system prompt, what's going to go into the user prompt? And if you go to your friends in CS, there is a whole literature trying to talk about different prompt engineering strategies to improve behavior. [02:24:04] So it then begs the question, do these additional auxiliary choices change the resulting measurements? [02:24:09] Um and this is not even bringing into the fold there is additional variation that arises from inherent randomness in model outputs and changes in model releases that are often hidden to users. [02:24:20] Um and some of the data editors at at economics journals actually put out an article last week flagging those concerns as well. [02:24:27] So I want to illustrate what happens. So I'm going to think about I'm going to take a data set um of congressional bills through 2018, about 10,000 of them that we randomly sampled. Um and for each bill we have a text description of it. We can also link that bill to the [02:24:41] party affiliation of the sponsor of the bill, the DW nominate, a sort of measure of their valence of the sponsor of the bill, and whether the bill originated in the House or the Senate. Um and some political scientists actually went through the very painful process of writing down an [02:24:54] extensive rubric, training a team of annotators for an entire semester. It's a really heroic effort to go and label all of the bills in Congress for their policy area. Is this a bill about defense spending, health spending, etc. [02:25:06] based on the text description? So for example, this could be a bill to revise the boundary of Crater Lake National Park in in Oregon. [02:25:13] Then what we did um in some work with Jens Ludwig and Sendhil Mullainathan did is we said, well, let's take their rubric and let's use various prompt engineering strategies from different models in OpenAI's GPT family to label each bill's policy area. And then let's [02:25:26] just run a bunch of regressions uh where we're going to plug in the language model label um on the left-hand side and then regress it on a bunch of these other economic covariates and just see what happens. [02:25:37] Uh if you do this, what you get is that across models and prompts uh your estimates can wildly vary both in terms of magnitude, significance, and even sign. Uh different models um will produce wildly different estimates, and it wasn't obvious to us, you know, X [02:25:52] ante when we first did this in 2024, whether there would be improvements um as models improved. So, a year later, we revisited it with GPT-5 and the the problem. [02:26:01] Let me do this in another example. There's a really wonderful piece of computational social science that studies uh speeches on immigration in the US Congress over about 140 years. [02:26:10] And what they wanted to do is they wanted to take each speech, they observe the text, they observe who was the speaker of the speech, what chamber they were speaking in, and what uh what uh and yeah, within the Senate or Congress, and also their partisan affiliation. And what they would like to do is they want [02:26:24] to take this to the text of that speech and label it. If this is a speech about immigration, is the speaker speaking about immigration uh in a pro-immigration manner, in a neutral manner, or in anti-immigration uh manner? And then see how that correlates with partisan identity over time. [02:26:38] So, you know, you could now revisit this exercise in 2026, use various prompt engineering strategies, take the rubric they wrote down, now incorporating models from OpenAI's family, Anthropic's families, Gemini, as well, to label the tone of each speech segment. And then [02:26:53] again, just plug those in on the left-hand side and see what happens. Uh and again, you see that across models and prompts, the estimated partisan tone gap differs enormously in magnitude and in sign. [02:27:06] So, what's the point? Given the rubric, there are still many important choices that an empirical researcher would have to make, which model, which prompt, and those choices, seemingly innocuous to us, can really move estimates in magnitude, significance, and even sign. [02:27:20] Um and there's a lot of work documenting this fact across computational social science. [02:27:24] So, you can look at this variation and have one answer. Well, why am I treating the model as a proxy? Let's stop doing that. The construct we're interested in is the model's output. The measurement process is then going to be defined by a particular model under a particular [02:27:39] prompt. Um you know, I think that's a complicated route you could go down uh because now then the S demand is going to be indexed by the choice of model and prompt. [02:27:48] Um are researchers using different models and prompts really studying different S demands? How is we How are we as a field supposed to manage the choice of GPT Claude versus Gemini, the choice of prompt? Are we supposed to revisit empirical work every single time [02:28:01] a new frontier model releases? Um I think that's a a set of challenging questions you'd have to answer if you wanted to go down this path. [02:28:09] Now, let me just flag there's been really interesting work over the last year where researchers are trying to build better harnesses, sort of the environment in which a language model sits to try to make them more robust to different prompt engineering strategy. [02:28:21] Uh this is the only example I know of people trying to do this at scale. Um and so this is an important avenue for us to try to really invest in. How do we build harnesses to reduce sensitivity? [02:28:32] But let me just bracket that off for now. So, we're not going to be thinking of defining the construct as the model, but again, maybe we still think there's not an existing measurement that we have access to. So, what would be an alternative view we could try to go down? [02:28:45] An alternative view we could try to go down would be to say, "Look, the construct we're studying is something that's inherently latent. There is no feasible measurement process that could recover it, whether human or machine. [02:28:56] Every single label we generate, whether by human, RA, expert, or AI, is just an imperfect proxy for this construct of interest." So, if you took this view seriously and then you know, went back to your conometrics classes, you'd say, "Well, [02:29:10] this sounds a lot like saying that we have potentially many repeated measurements of this same latent construct even though we never directly observe it." So, could we then try to take classic ideas from the measurement error literature to make progress? [02:29:24] So, you know, let's just do a reminder, what would be one classical identification strategy from the measurement error literature? [02:29:30] So, now I'm going to let little K index different model and prompt combinations, where R I K is going to be the measurement generated by that particular model prompt combination. [02:29:40] And let's imagine this construct we're interested is binary, just zero one. And the thing I want to learn about is just its base rate. How often is the construct equal to one? [02:29:50] So, the key assumption in a lot of this work is going to be to assume that if I were to condition on this latent construct, then these measurements are independent of each other. If I was in a classical measurement world, I have two measurements, they're independent of each other, I could use one as an [02:30:03] instrument for another. That idea really generalizes. [02:30:07] Um and in particular, the intuition you can have in the back of your mind, but you could spend time and prove it more carefully, is that if I was willing to make this assumption, what I get to learn from the data are two to the K minus one probabilities. There are two [02:30:21] to the K plus one unknowns that I'm trying to infer, and as long as K is bigger than or equal to three, I have enough equations relative to unknowns, so that maybe I could solve this problem. [02:30:31] Um and indeed that intuition is right, if you're willing to make some auxiliary assumptions, and in particular those auxiliary assumptions would be that these measurements are informative about the latent construct of interest, meaning the probability of a true positive is larger than the probability [02:30:44] of a false positive. And you can generalize this in a lot of ways. There's a lot of fantastic work in measurement error trying to take this idea and run with it as far as you can. [02:30:53] So, what if we try to take this idea to large language models? [02:30:57] Well, now if we take this idea to large language models, the measurement error assumptions we're writing down are properties of measurements. Here the measurements are outputs of large language models. So, we're really writing down assumptions about the behavior of frontier AI. So, we have to [02:31:11] ask ourselves the question, which of these assumptions are credible? [02:31:14] So, in this simple case, the additional informativeness assumption feels very mild. Presumably, if you were a researcher willing to use a language model to produce measurements of some construct, you probably were willing to believe this up front. [02:31:28] The problem is going to come in this conditional independence assumption. Is that really going to be credible for large language models? It would say that co-movements in the outputs produced by these language models, different prompts, is only going to be driven by this latent construct. Errors therefore [02:31:42] must be unrelated across models and prompts conditional on the latent construct. Why might we worry about this? Of course, there is substantial overlap across frontier AI labs. There is substantial overlap in the sorts of data they pre-train on. There may be [02:31:55] differences in what is the data they use for uh you know, fine-tuning on human feedback, fine-tuning on verifiable rewards, but they're presumably drawing those from similar domains. And also, there are shared architectures across these models as well. And this is you know, not just something that we're [02:32:09] worried about. There's a very active literature in computer science under the lens of algorithmic monoculture that tries to think about this problem uh more seriously. [02:32:18] But ultimately, you know, we can talk about this and why we might be worried. [02:32:21] This is ultimately an empirical question about the behavior of these models. And where do we understand or how do we learn about the behavior of these models in practice? [02:32:30] So, if you are uh too online like myself, you've probably seen every time a new language model is released, uh widespread publicity around benchmark evaluations. The real quantitative basis of our understanding about model [02:32:43] performance comes from benchmarks. And you can think of these as looking at the performance of 3 years ago, GPT-3.5 on standardized exams. Then people then move to more domain-specific exams. As those got saturated, we moved to [02:32:57] benchmarks that try to test more general intelligence or coding benchmarks, math problems, etc. So, in other words, we evaluate performance by asking these models questions and evaluating the outputs. [02:33:07] So, we could try to do this to understand this measurement error assumption. Um rather than just talking about why errors might be correlated, let's actually go and look. So, as I mentioned, there's this active literature in computer science under the umbrella of algorithmic monoculture, and [02:33:21] they went and did this exercise on existing benchmarks used to evaluate frontier AI. It's a really wonderful paper by Nico Guard who is at Cornell in his lab. So, what did they do? Is they looked across pairs of models on different benchmarks, and then they [02:33:36] reported how often two models gave the same wrong answer conditional on being wrong. And they did this across a large set of models that are publicly available on hugging face and then a large set of benchmarks that are publicly available as well. [02:33:49] And what they found is that the wrong answer agreement across different models and prompts far exceeds what you would expect from random chance. And what's even more interesting is that this correlated errors problem is more is more problematic and you know, more [02:34:03] common on better performing models. It seems like with new model releases outputs are becoming more and more correlated. And if you go to the paper, they have some really you know, wonderful little heat maps that you can kind of stare at and try to understand and look at the correlation structure. [02:34:16] But the point is it's not blue. There's a lot of yellow here. Things are very correlated. [02:34:22] But again, you know, you may be skeptical and say, "Look, I don't care. [02:34:26] This is evidence about AI benchmarks. We're talking about economics. What about the measurement problems that we as economists are interested in studying?" So, let's now revisit two running example or two running examples and assess whether measurements from different models and prompts are in fact conditionally independent. So, let's go [02:34:41] to this congressional legislation example. Let's go to this data set on immigration speeches in Congress and ask conditional on what the authors treated as their existing measurements, would model errors be correlated with each other? So, you can implement that [02:34:56] in a more precise way. Let's generate labels from a bunch of different model and prompt combinations. For each pair of model and prompt combinations, let's calculate their joint distribution and compare that to what we would expect to see if in fact they were independent [02:35:10] conditional on the construct. And then we can summarize that in a test statistic and an associated critical value. So, what I'm just going to show you next is what is the ratio of the observed test statistic relative to its 95% critical value, which is just a [02:35:23] simple summary of getting a sense of how strong is the evidence against things being independent of each other conditional on the construct. [02:35:32] Um and what you find is in the congressional legislation example, if you look within a model across different prompts, perhaps unsurprisingly each prompt pair looks to be significantly correlated conditional on the ground truth. [02:35:45] It's not surprising. This is the same model prompted in a different way. We would expect there probably to be a lot of correlation. [02:35:50] Um if you do the same thing in the immigration speeches, you find the same pattern, but interestingly on trying to label sentiment about immigration speeches, there is more correlation conditional on the ground truth in errors across language models across prompts within a model. [02:36:04] And then you can also do this across different models. So, now each different row is a different model, whether GPT or Anthropic. And interestingly again, you see while there is less correlation across model families, things are still [02:36:18] severely correlated in general. So, what is the takeaway? On benchmarks and in economic applications, it appears that errors are strongly correlated across models and prompts. So, if we were to go to the measurement error [02:36:32] literature and said, well, let me take this idea, let me treat language model labels across models and prompts as if they are conditionally independent, that would not be credible. [02:36:40] So, identification results that rely on it really should not be naively applied to LLM generated measurements. [02:36:46] But at the same time, you know, that seems like a very very pessimistic take away from this problem. The errors are correlated, but they're not arbitrarily correlated either. We saw differences when we looked across models as opposed to across prompts within a model. [02:37:00] Taking a step back, to me the cleanest framework for trying to incorporate AI outputs into empirical research remains through the validation sample. [02:37:09] But I think that framework is often properly is often misunderstood. [02:37:12] Properly understood, you should think of this as its goal is to try to automate and scale our measurements at our best. [02:37:19] And best is not the compromise we settled with before AI. It's the measurement process you would actually defend, which is a documented, adjudicated rubric decided upon and applied by you the researcher, not the undergrad RA that you were able to [02:37:33] cheaply hire. And the reason for that, the reason why this framework is powerful is that our best measurements no longer have to cover the entire corpus. A small expert labeled sample may be affordable, and now you can use AI to scale that measurement across the corpus of [02:37:47] interest. So, in other words, and just coming back to a point that Melissa emphasized, really what we now have in the age of AI is that the scarce input for us is constructing high-quality measurements in the first place, and that's what we as researchers must be [02:38:02] investing in. And as Melissa alluded to, we've been here before. When running regressions became cheap, the attention of econometrics, the attention of empirical research wasn't about how do I invert X'X inverse, it would became on how do we defend and [02:38:16] find credible identification strategies out in the world, and then let it rip from there. And a similar thing has to happen here, invest in high-quality measurements, let it rip with the validation sample framework. [02:38:26] So, the second takeaway is that once you step out of this framework, you're now in a world of making assumptions about measurement error. And making assumptions about measurement error are in fact making claims about the performance of AI and ML models. And this is a place where economics can do a [02:38:41] lot of work and sort of take a step out of our comfort zone and start engaging with work in AI, where there's already an entire field that tries to study frontier model behavior through constructing benchmarks. And what I hopefully illustrated in that example was we have a huge literature in [02:38:55] econometrics that thinks about measurement error assumptions through these conditional independences. We could have assumed it ex ante, but if we went and checked it empirically, we'd find it falsified. But perhaps this exercise could provide some external calibration. And you can go to, you [02:39:08] know, many long survey articles about deep ideas in measurement error and ask, "What would be the analogs of these strategies for AI models? How would we benchmark? How would we evaluate whether they're credible or not?" Um and so, you know, I just want to end [02:39:21] on a point that I think uh this, you know, section of of the the talk is that, you know, while there's a ton of work on benchmarking in CS, these are not the benchmarks that we as researchers need. They focus on the wrong tasks. They're thinking about, you know, math exams, coding, not [02:39:35] measurement problems. And, you know, we're generally bad at generalizing performance to new problems. And furthermore, they often focus on the wrong properties, thinking about just overall accuracy, which can be a misleading guide to what is going to be the bias for downstream estimates. So, [02:39:49] developing and working on benchmarks for measurement in economics and the social sciences more generally is a really essential activity, one where we invest in constructing measurements and then asking, "How well do models reproduce or [02:40:01] how do they reproduce those?" So, so far, what we tried to do here was to depart from this validation sample framework in one way, where we ask the question, "Suppose there is no existing measurement that we'd like to scale." And we talked about how measurement [02:40:15] error assumptions are going to be claims about frontier model behavior. And benchmarks are how we could in principle discipline those assumptions. [02:40:22] What I next want to talk about is another deviation from the validation sample framework. And that's going to be one in which we're not disputing the facts. There is an existing measurement we would care about. But that measurement exists in some other sample. [02:40:34] So, one idea you can have in the back of your mind would be ground truth, the object we're interested in is some consumption measure that we would collect if we got a household in some remote village to sit down for many, many hours and ask them detailed questions about everything they ate in [02:40:49] the last week. We like that as our measure of consumption. The problem is is that was only collected in a census or a survey in some other region in some other time period. And this often most most often arises in remote sensing applications with satellite imagery or [02:41:03] uses of digital trace data. So, think about the trace data associated with a cell phone usage. [02:41:09] And in this case, the researcher is going to have to somehow combine the sample that they have of interest with this auxiliary sample, and identification is going to really turn on what is assumed to be stable across them. [02:41:20] So, in this case, you know, the constraint for the researcher and what AI and ML may hope to alleviate is cost and coverage in in collecting outcomes or covariates that we're interested in studying. [02:41:32] There's been a ton of work about how unstructured data in the form of satellite imagery and digital traces, this mobile phone data, could help us scale existing measurements. And the reason being that satellite images, mobile phone data are collected at very [02:41:45] high frequencies, at very small granular geographic level geographic areas. And they're already collected by other companies, other space organizations. [02:41:56] So, how can we take advantage of these unstructured data to lower measurement costs in experiments and observational studies? Most often, we see this in environment and development economics. [02:42:07] So, what are we going to think about? So, again, researcher collects unstructured data XI. Now, you can think about that as satellite imagery or this digital trace data for each unit. Each unit is associated with some construct of interest V star. Think of this as our [02:42:22] survey measure or census measure of construction of consumption or perhaps some local environmental outcome. And then we may have access to a prediction RI of this construct that is built from the unstructured data. So, a prediction [02:42:35] of local consumption from satellite imagery, a prediction of local pollution from satellite imagery as well. And there's been a lot of work while this, you know, this this is a very popular empirical idea, there's been really wonderful work over the last 3 to 4 years documenting that while these [02:42:49] remotely sensed predictions can be very accurate, they're still imperfect. There still is error and there's error that's correlated across space and time. [02:42:57] So, you could take the ideas from Melissa's lecture and apply this again here if you, the researcher, have a particular sample of locations in space, villages, collect your remotely sensed proxy on all observations, randomly pick a subset [02:43:12] of villages or households to actually go and run the survey or collect the environmental measurements, and then again go through debias. [02:43:21] But of course, collecting a validation sample, you know, there there there's cost involved here. Um and oftentimes the labels we care about can exist elsewhere. There may be a government environmental survey, a census from earlier survey rounds, um but those are often collected somewhere else. Um and [02:43:35] so the idea would be could we somehow use those, treat those like a validation sample, and I'll call them an auxiliary sample, where we observe ground truth, our predictions, and maybe additional covariates, but in the sample that we would like to run our experiment, in the sample we would like [02:43:50] to run our observational study, we don't have that ground truth. We just have the remotely sensed prediction, covariates, and maybe some other treatments of interest. [02:43:57] So, can we use these existing labels to correct for the errors in RI? This is now fundamentally a data combination problem. [02:44:06] So, if, let me just push back on one thing you thought you could have, well, I've heard ML AI really good at predicting, maybe these predictions are going to perform well across time and space. [02:44:17] Um if we had that view, what would we expect? Um well, you know, as I mentioned benchmarks earlier, there's a really wonderful recent paper by uh Jonathan Proctor and a team of folks in environmental and development economics, where they went they did the hard work of constructing a benchmark, where they [02:44:31] said, let's take um some pre-trained embeddings of satellite imagery called the Mosaic framework. That is itself a wonderful piece of computer science output. It is a publicly available API where anyone can go and pull embeddings of satellite imagery across the world. [02:44:46] It's amazing, easy to use. And they said for a bun for about 115 socioeconomic and environmental variables, let's take this pre-trained embedding of satellite imagery, predict those variables using satellites, and then ask, how does the performance of those predictions vary [02:45:01] across time and across space? Um and what you find is that as you uh move further and further away from the training sample in geography across every single class of outcomes they look at, uh there are large reductions uh in [02:45:15] predicted performance. So, if you're trying to merge samples across space, you should be worried about instability in the overall accuracy of predictions. [02:45:24] You know, you can run the same exercise We can run the same exercises ourselves. [02:45:28] I just wanted to do this in the lecture to illustrate how easy this is. This was about a week of work in collaboration with some CodeX agents. Again, using stuff that all you can all pull from publicly available APIs for mosaics, and then there are other satellite [02:45:42] embeddings as well called Alpha Earth. So, just going to show you another example where we're going to try to predict a outcome index about living standards, um which is a 14-component index of housing variables, infrastructure, assets, and financial [02:45:56] access uh from the the DHS survey. Train that on a bunch of geographic clusters in Kenya, and then ask how well the predictors are going to transfer to different parts of Kenya or Ethiopia or a different point in time, and then do the same thing uh to predict forest [02:46:11] cover in Uganda, which is the the setting um where Seema Jayachandran and co-authors had a really wonderful early-stage use of remote sensing in an experiment on deforestation. [02:46:21] Um and again, what you find is if you try to transfer these predictors for mosaics and then a more updated or more recent uh satellite embeddings Alpha Earth, you see a lot of transfer error both across time and space, and that's true both for the living standards [02:46:34] outcome and for the forest cover outcome. [02:46:38] So, we can't just take these objects and assume, you know, on average they're going to predict as well. We're going to have to introduce some assumptions here. [02:46:45] So, let me just add a bit more notation in the last 5 minutes. Bear with me. [02:46:48] We'll get through this together. Um so, we're going to think about a researcher that has two samples. I'm going to call S equals E, their empirical sample. That's the sample where they're only going to observe W and the remotely sensed predictor. And an auxiliary sample where they're going [02:47:02] to observe both the remotely sensed predictor and their existing measurements. [02:47:07] So, the target is going to be some parameter in the experiment in empirical sample, say the base rate of this outcome or perhaps we have some treatments and we'd like to learn the treatment effect of D on this outcome of interest. How could we combine these two together? [02:47:20] So, one way we might try to approach this data combination problem uh is an assumption that I'm going to call here outcome stability, which in two at a high level, what is it saying? Is it saying that our predictions of V star given R I and then possibly additional [02:47:34] covariates, that prediction transfers across the samples. So, conditional on seeing the same satellite image and having the same covariates, the distribution of the outcome of interest would be the same. [02:47:46] If you believe that assumption, you could uh you know, take yourself back 1 year and you would find yourself in a framework that looks a lot like the surrogacy framework and we could apply identification results from there to learn objects of interest in the [02:47:59] empirical sample. And it's important to sort of think about what is this saying in terms of properties of the predictor? Uh outcome stability is again sort of saying that the predictor has in machine learning has the same calibration curve in both [02:48:12] samples. Um and I'm going to draw an an analogy here to folks that are familiar with this literature. This is a a classic notion of fairness in computer science. What it means for a predictor to be fair would be that the group membership uh does not affect [02:48:26] calibration. Um and I'm going to come back to that. That's going to be a thing that we open, we'll close in a little bit. [02:48:33] An alternative assumption you could make which would be different than outcome stability is what I'm going to call measurement stability. Measurement stability would say that conditional on the latent out the outcome of interest and additional covariates, the measurements, the predictors themselves [02:48:46] RI are going to be the same across those two samples. [02:48:50] So, as an example that you could keep in the back of your head, imagine we were trying to infer crop burning from satellite imagery, the underlying crop burning process is what's going to produce the visual signal in satellite imagery, and our idea might be that crop [02:49:04] burning looks the same in satellite images whether that is measured in one part of Malawi or another part of Malawi. And so, I'm willing to believe that that process is the same across those two samples. [02:49:16] Uh if you were to believe this assumption um in some recent work that I have with Rahul Singh and Davide Viviano, you can work through and still identify quantities of interest in the empirical sample uh of interest. Um it's [02:49:28] going to look like uh sort of a an IV type problem, conditional moment restriction problem, but I'll punt on all those details for now. [02:49:36] And again, I just want to pause and try to interpret what this restriction is actually saying. If we're in a world where the outcome we're trying to measure is binary, and we had some classifier based on satellite imagery, so RI is also binary, measurement [02:49:50] stability would say that the errors of this classifier are going to be conditional on the latent outcome of interest would be the same across those two samples. My predictor should have the same true positive rate and false positive rate uh in both samples. And again, going to open another bracket, if [02:50:04] you went to work on algorithmic fairness, uh this is a diff another notion of what it might mean for predictors to behave the same across two samples. [02:50:13] So, what's the relationship between the two of them? Um well, if you were to suppose that outcome stability holds, you can go through Bayes' rule, a nice little first-year exercise to ask what is its implication for the measurement channel, and you would see that if outcome stability holds, then [02:50:28] measurement stability can only hold if there is a nice cancellation property of the base rates in the outcome and the unstructured data across the two samples. [02:50:37] And if you went in the reverse direction and asked, suppose measurement stability holds, what would it imply for an outcome stability? Again, outcome stability would only hold if there is nice cancellation in particular terms that depend on the base rates across [02:50:50] those two samples. So, it would seem like having both of them hold at the same time is a very knife-edge case. I would need some specific properties about how outcomes are varying across these two samples. [02:51:02] Um but differing base rates are exactly what you would expect. The auxiliary sample is often collected somewhere else at some other point in time. A government forest survey run in a different part of Uganda might have a different level of forest cover than the sample you're studying. Mobile phone [02:51:17] carrier data from Malawi matched to a particular census year may have different income or consumption patterns than the experiment you're running, say, post-COVID. [02:51:27] Um that's actually an idea you can make very formal um and actually show that outcome stability and measurement stability can both only can hold only if the base rates agree across the samples or the base rates differ, but you're in this world where your predictor is [02:51:41] perfect. There's no error. Neither of those cases are very reasonable in practice. You would expect base rates to differ across samples if you're doing data combination. The benchmarks we just looked at told you that predictors are imperfect. Um and [02:51:56] while that may seem like a a mysterious result at this, uh you know, very fast rate, uh this is just a restatement of some very classic canonical impossibility results in algorithmic fairness. Basically, the two stability assumptions we wrote down are two different notions of what it means for [02:52:10] an algorithm to behave the same across subpopulations. Um and there's work going back now about a decade ago showing that those two things can't hold in general at the same time. [02:52:20] So, what's the upshot? You know, these two stability assumptions are very powerful. If you apply them, you can identify the stuff you're interested in, but they're alternatives from one another in general and in practice. So, you got to pick and you got to defend the choice. [02:52:31] Um I'm going to go through this part a little quickly. [02:52:35] Uh, I'll do my sort of, you know, econometrics police work. If you choose the wrong assumption, you're going to get wrong answers. So, choose the right assumption, think hard. [02:52:43] Uh, but if you get the correct assumption right, uh, there's, you know, really high returns to this. So, let me just illustrate very briefly a semi-synthetic exercise based on this deforestation experiment idea using data from the DRC, Brazil, and Indonesia, um, where we're going to try to predict [02:52:57] forest cover, but we're going to set this up so measurement stability holds, and I want to just illustrate the value of exploiting an assumption when it's correct. [02:53:05] Um, the blue line is going to show that, uh, if measurement stability holds and you use an estimator that exploits it, you're going to get estimates that are approximately unbiased across a variety of data generating processes. If you use an assumption that does not exploit measure estimator that does not exploit [02:53:19] measurement, uh, measurement stability, you can get wildly incorrect estimates. [02:53:24] Furthermore, if you exploit the information available in the auxiliary sample correctly, uh, you can really substantially tighten standard errors by taking the assumption seriously and writing down an estimator that exploits it. So, the consequences of getting the [02:53:37] assumption wrong are bad, but if you get it right, there's really huge upside here in terms of getting unbiased estimates and meaningful improvements in standard errors on the order of, uh, 20% relative to some alternatives and even [02:53:49] 40% relative to some alternatives. So, uh, you can do this my many other empirical settings. Um, I'll refer you to this, uh, this paper with Rahul Singh and David Dhavale where we do this in two empirical illustrations. Um, but [02:54:02] really that's focusing on measurement stability. There are large returns in general to exploiting the right stability assumption and doing data combination with ML and AI generated measurements. Uh, so, how should researchers choose? Um, we're getting close to, actually we're at over 6:00 [02:54:17] p.m. So, I will punt on that and say there's been a lot of work over the last year to try to give people guidance. ## Q&A (02:54:22 – 02:56:26) [02:54:22] It's in the slides, it's in the papers. Reach out if you have any questions. [02:54:25] Um, and let me just sort of try to wrap up here on what we tried to accomplish across these lectures. You know, for me, AI and machine learning are really accelerating the process the discovery process across the sciences. I think it's a foundational and important [02:54:39] question that we must wrestle with. What would this look like in economics? Our goal in this lecture was to try to answer one provide one answer to that question, which is that machine learning and AI allow us to tackle a measurement problem that we've always faced. The first part is that we collect limited [02:54:53] structured measurements from the settings we study, and the second is in settings we study collecting high-quality measurements is very costly, and both of these limit the questions that we can study at scale. [02:55:04] Um, and what we tried to argue for you is that uh there are some real opportunities for ML and AI to tackle both those problems. So, the first problem, we collect limited structured measurements, AI provides us with the opportunity to expand the aperture. We are building ML and AI tools to do [02:55:19] hypothesis generation from unstructured data. That can give us new ideas to wrestle with using our standard empirical toolkit. And the second problem is that measurement is fundamentally costly in the social sciences, but we can use ML and AI to scale the measurements that we have and [02:55:32] we would already defend. And there's really good work thinking about this problem where you have access to a random validation sample, uh and we're starting to see work in the space where there is no random validation sample and we have to deal with data combination. [02:55:44] You know, my last slide that I do in my class, I'll do it for you guys, is I think over the last 25 years or the crowning achievement of statistics and econometrics has been the causal inference revolution. Gave us new ways of analyzing data that unlock new ways to see the world and open new questions [02:55:59] that we couldn't study before. And I think we are now building the same sorts of toolkits to use machine learning and AI as tools for measurement. And measuring better allows us to see the world in new ways and open up new questions that we couldn't study before. [02:56:11] Um, so I hope that, you know, we as a field take up this challenge. I hope that you go out and measure interesting stuff. And please, if you have any questions, I'll I'll offer my my inbox. [02:56:21] Please reach Love to chat more. >> [applause]