Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # Measurement with Validation Samples: A Missing-Data Perspective, Melissa Dell, Harvard University and NBER Authors: Discussant: None Video: https://www.youtube.com/watch?v=ydYXe2FxErk&t=5682s ## Talk (01:34:42 – 02:15:32) [01:34:42] >> All right. Um, so I'm going to be spending the slot talking briefly about construct validity and then talking about observational validity. Um, and you know, one of the things I've been told is that oh, I'm not on are the [01:34:56] sound people? Okay. Ken, is this working? [01:35:02] All right. Uh, so as I was saying, I'm going to be uh, spending the slot talking about construct uh, validity very briefly as well as observational validity, which is a much more developed literature. Um, and you know, so one thing I was told is that, you know, the [01:35:16] main reason uh, people might uh, choose to view the lectures or attend is they want to know, you know, how should uh, how should I compute my standard errors? [01:35:24] How do I run a differences-in-differences specification uh, with AI predicted variables? And we're going to get to that um, in this um, 40 minutes and I will introduce an R package for doing that. Um, but before I get to those very practical questions, [01:35:38] I'm going to spend just a few minutes um, talking about kind of a much more open area uh, because I think it's important. Um, and you know, I don't have an R package or Python package or for of for this. Um, I don't even know [01:35:52] exactly the right way to approach it, um, but I'm going to briefly talk about it, uh, because I think, um, that it's a it's a huge open area and very important. Um, and so if you recall, um, from my first lecture, construct validity addresses how a theoretically [01:36:07] meaningful concept should be operationalized. There's conceptual uncertainty, and this is important because economists are increasingly measuring novel, latent, high-dimensional concepts that are not [01:36:19] directly verifiable. Um, and in the age of AI, we can potentially create many different measures cheaply, um, and that's great because we can explore measurement. We can also create more evidence for validating those measures [01:36:33] cheaply, but at the same time, it raises this concern that I introduced with the economic policy uncertainty index, um, that you can create many different measures of the same underlying reality. [01:36:47] Um, and so which of those are valid and which of those aren't? We need to be able to answer that. I think it also raises another set of issues, which I don't have anything to say about, but that I just want to flag, which is that construct validity is use contingent. [01:37:01] You might have a measure validated as a passive description that could lose validity once it's embedded in decisions and agents re-optimize. And because AI makes it cheap to deploy it, um, uh, [01:37:13] constructs in real time at scale, it could exacerbate this problem. [01:37:18] Essentially, kind of an analog for construct definition of Goodhart's law. [01:37:23] Uh, but I want to focus on construct validity as partial identification. Um, and as I said, this is a pretty open area. I mean, people have thought a lot about this kind of philosophically, but [01:37:36] in terms of having an econometric framework, there's, um, you know, much, uh, much less literature. And so the way that it seemed most natural to me, uh, to approach this as a first pass was to [01:37:49] take the nomological network of Cronbach and Meehl um and apply that as a partial identification problem. And so, remember the idea here is that if we cannot directly validate a measure because it's high dimensional, it's a latent [01:38:03] construct, it's not directly observable, um we want to be able to kind of validate it uh with multiple restrictions to see which measures are valid and which it doesn't. And so, the researcher has to specify the candidate [01:38:17] measure. The researcher has to specify a set of moment restrictions that theory or the existing literature or knowledge of the context suggests that those measures should meet. Um both require justification. [01:38:29] Um of course, this isn't giving a causal account of measurement, um but it is helping to discipline it. Um and so, I'm going to let X denote high dimensional reality, um and a researcher can considers a menu of different [01:38:44] candidate uh measurement functions, uh with each candidate mapping the same underlying reality into a measure. Um both the measures and the downstream estimates must be made comparable before we can do this exercise, which you know, [01:38:58] you might have an indicator and a scalar measure, you need to put them on the same scale, the downstream estimate uh beta also needs to be defined consistently so that you can compare measures. Um which is relatively straightforward. Um [01:39:11] now, I'm going to define auxiliary observables um that are implicated, you know, by theory, by our knowledge of the context. [01:39:19] Again, we might have related quantity Z, we want the measure to correlate with them. We might have distinct quantities, we don't want the measure to capture that because that's something separate that we want to hold constant. Uh we might have groups that should differ on the construct. So, for instance, if we [01:39:34] want to measure partisan speech, we want Republicans and Democrats to differ on that dimension of speech. Um and so with these variables we can create moment restrictions. Convergent validity says [01:39:49] that um, the measure needs to be correlated with the related quantities. Discriminant validity, it needs to be distinct from the things that that we want it to be distinct from, um, etc. And you can make other moment [01:40:01] restrictions as well. And so we're going to collect the threshold and express these restrictions as a population moment inequality. And given a fixed uh, candidate menu, we can determine the admissible subset. These [01:40:16] are the ones that meet all of the moment restrictions. Uh, this is a population object. It encodes conceptual uncertainty, um, you know, separate from sampling uncertainty that comes from estimating it. [01:40:28] Um, for each candidate measure we also want to report slack in the restrictions because we want to understand which restrictions drive admissibility. Are there certain restrictions that are really important in determining if this measure is valid or not? Um, would [01:40:43] conclusions change if we used alternative thresholds, which hopefully are kind of um, are are supported by um, benchmarks, prior evidence, etc. Um, are conclusions robust or do they appear [01:40:55] relatively fragile? What we ultimately care about is some construct identified set, um, such as a regression coefficient, a treatment effect, a structural parameter, [01:41:07] um, and given this we can define the set of estimate target estimates, um, that um, we obtain through taking the admissible measures and using them in our downstream analysis. [01:41:21] If we have a scalar target, um, our goal is to estimate a confidence region. You know, so we are interested in the impact of economic policy uncertainty on employment and we have different measures of economic policy uncertainty [01:41:35] that are consistent, um, with our moment conditions and now we want to see well how much does the regression coefficient of interest change if we use these different measures that are consistent with the theory [01:41:49] and so there's uncertainty in beta hat for each candidate measure and there's also uncertainty in the estimated moment inequalities that determine which measures are admissible which measures are consistent with theory and these can [01:42:04] be coupled and we're not going to talk about estimation today but existing partial identification and endpoint inference methods can be combined to estimate a valid if conservative confidence region as I said this is an important you know [01:42:18] question um and I think it's something that I'd like to encourage more research on you know hopefully there will be a package soon that automates this and makes thinking about this a bit easier but I think that there's also a lot of scope for [01:42:31] methodological advancements and I think with AI we just can't avoid this question all right um and so I'm going to go ahead now and move away [01:42:45] from construct validity and talk about observational validity which is a quite a contrast because the literature is actually really well developed for observational validity [01:42:59] okay so again to recap what I mean by observation with observation you already have a specified construct and you want to take that and implement that at scale deep neural networks are the state of [01:43:14] the art tool for large scale feature extraction from unstructured high dimensional data and are increasingly being used by economists by individual researchers to construct data on a large scale for empirical research [01:43:29] importantly though we cannot assume that neural networks will generically produce unbiased predictions in finite samples. [01:43:37] Systematic biases can arise from the network architecture, the distribution of the training data, other implementation details. [01:43:44] Um, we're taking non-linear transformations in the neural network. [01:43:47] We're applying it to binary or multi-class classification. This all violates classical measurement error assumptions. And so, to the extent um that the AI makes mistakes, and if you use AI, you see that it always says at the bottom of the screen, it may make [01:44:02] mistakes, and indeed it does. Um, we we can't assume that those are classical. [01:44:08] And so, typically, you know, we oftentimes think of structured data as proxies. Um, and this is what was feasible to do historically. Um, you know, it was very [01:44:20] costly to uh to create data by hand, and oftentimes you had to make do with some existing measure that existed. But, in a world of expensive of inexpensive AI, you know, such as commercial large language models, if we treat observation [01:44:34] as the construction of proxies, this poses quite serious challenges for the integrity of empirical work. [01:44:42] Okay, so the first challenge is that sh- choosing different models, different training data distributions, different implementation details that are unrelated to the construct definition itself can produce different predictions. Um, [01:44:57] indeed, that this happens, um and it can lead to bias. Um, it can make it difficult to interpret the results, and it leads to post-selection inference concerns. [01:45:08] We could try to ex- to address this with sensitivity checks, but as I argued in the last lecture, they're not going to fully address this concern because biases can be systematically correlated across models. [01:45:21] Moreover, economists frequently use proprietary um large language models, proprietary neural networks, and this raises reproducibility concerns. Um so, suppose you use Opus 4A for your paper, um but [01:45:35] by the time the paper is published years later, that has long since been deprecated, and you go to use kind of the most um the most recent model, and you get a different answer. That that's a problem. [01:45:48] Um but at the same time, these models are so powerful, like we wouldn't want to not use them because of that issue. [01:45:55] And then finally, it's typically possible to improve neural network predictions given you have a well-defined kind of um objective through costly investments. Um you could, you know, train on more data, train on better data, uh for longer, you [01:46:08] know, have a customized model instead of using something off the shelf, but it's unclear how to assess which investments are worthwhile if you don't have a principled way to account for how prediction errors uh from your neural network affect the estimation of [01:46:23] whatever parameters are of interest for your research question. [01:46:28] So, in order to address these challenges, it requires a shift in perspective, um moving from informally choosing proxies that are plausible to first of all specifying the measurement aim. In other words, we need a [01:46:43] well-defined construct, and then validating observation. How well is that aim defined and implemented at scale? [01:46:52] And so, statistically, what we would like to have is a broadly applicable framework that can correct the biases in observation that come from imputing data with a neural network, um yielding target parameters that are robust to the [01:47:07] choice of measurement instrument. So, what does that mean? It means, you know, that if I use um Opus, if I use GPT 5 4, you know, if I use an open-source model, I will get, you know, up to up to error, [01:47:21] um statistical error, I will get the same answer. And we would also like better measurements to improve precision, which clarifies the cost-benefit trade-off of making investments in better measurement. [01:47:34] You know, unfortunately, just such a framework exists, and it's existed for 50 years, um, you know, more or less, um, which is, um, the missing at random, uh, mechanism by Rubin. [01:47:47] Um, and so you might be saying, "I thought we were talking about inference with AI predictions. Why are you talking about missing data?" Uh, well, observation inference with AI predictions, right, is framed in this way because we typically [01:48:01] lack the low-dimensional summaries needed for statistical analyses, uh, when we're using high-dimensional unstructured data. You know, so you have millions of newspaper articles. Those newspaper articles do not come with a [01:48:14] label that says, "This newspaper article discusses economic policy uncertainty, yes or no." That's missing data that you need to impute. And so that's why this is a missing data problem. Um, and it turns out that once we frame inference [01:48:28] with unstructured data as inference with missing structured data, there's a whole kind of large literatures, a whole range of tools that we can apply. Um, and in particular today, I'm going to talk [01:48:42] about a framework called missing at random structured data. [01:48:46] Um, this is joint work with Jake Carlson, who's an econ econometrician. [01:48:50] He's on the job market this year and doing lots of fantastic work kind of in this space of AI and inference. [01:48:56] Um, and this tries to make kind of a long-standing literature applicable to applied economics. [01:49:06] Okay. So the core [clears throat] idea of missing data problems is that we're going to use a validation sample to estimate the bias in the imputed structured data and adjust our estimates accordingly. [01:49:19] Um validation data come from some implementable construct definition that's specified by the researcher and treated as given. It may be a silly way to measure that, you know, particular concept, but it's it's stated [01:49:34] explicitly, it's taken as given, um and you take that construct and you apply it through some non-scalable process, you know, that could be everything ranging from, you know, ground station temperature measurements um in the case [01:49:47] of remote sensing data uh to export annotations in the case of tax. Uh I have a rubric for economic policy uncertainty, does this article discuss it? Yes or no. [01:49:57] Um in the key assumption in um these frameworks is missing at random. [01:50:03] And so, what that says is after adjusting for observables, your annotated and unannotated data are comparable in their ground truth values. [01:50:12] What do I mean when I use the term ground truth? Importantly, I'm using that in the machine learning sense, it doesn't mean fundamental truth. [01:50:18] Remember, the goal of measurement is not to uncover some fundamental truth, instead it just means the label that it would have if you applied your kind of well-stated criteria uh to that particular observation. [01:50:30] This means that there's no unaccounted confounders that determine whether an instance of unstructured data is annotated or not. [01:50:41] Okay. So, semi-parametric inference lets the data speak for itself as much as possible, um and why is this so useful when we're doing inference with AI predictions? [01:50:52] Well, the missing at random assumption makes um you know, minimal assumptions about the neural network itself. And so, by having missing at random, we don't have to make um assumption, you know, strong assumptions about the neural [01:51:07] network and how its biases shift, which we don't want to do because they shift in really complex ways with the inputs, with kind of what sort of unstructured data we feed into them. And this has already been referenced several times. [01:51:20] People have difficulty predicting how this domain shift affects performance. [01:51:24] And so if we can avoid it, we really really don't want to have to make an assumption of like, you know, well, you know, this this model with satellite data was estimated in the United States and when we go and estimate, you know, [01:51:37] how it predicts crop type in South Sudan instead the errors are going to be the same, right? Because we would expect the biases to shift in complex ways and there's a large literature that shows that they do. And so by assuming kind of missing at random, we're able to make [01:51:52] minimal assumptions about the neural network. Of course, if we don't have missing at random, you're going to have to impose stronger restrictions and that's what the final lecture is going to be about because of course it's not always feasible to impose this missing at random [01:52:06] assumption on the annotated data. Okay. [01:52:11] So this idea of debiasing, kind of essentially collecting a validation sample and adjusting for the measurement error, is not a new idea. Okay, there's large literatures on measurement error, [01:52:24] semi-parametric missing data, causal inference, valid inference with black box AI predictions. Actually, this problem is very closely related to causal inference because causal inference is also a missing data problem where it's your, you know, counterfactual treatment outcomes that [01:52:38] are missing. And so the idea of debiasing is not new, it's very well established. Um, what the MARS framework does is to try to make this as accessible as possible to applied economists who are working [01:52:53] with AI predictions. We try to unify insights from these literatures into one applied framework, identify estimators that are both unbiased and efficient, and finally extend debiasing to common settings with limited guidance, including aggregated [01:53:07] missing data. Um and we do have an R package that will allow you to to produce um all the estimates um that that you'll see me present today. [01:53:17] Um And so hopefully it makes it straightforward um for for people to use this. [01:53:24] All right. I want to give a motivating example. Um this comes from uh Baker, Bloom, and Davis uh in their paper on measuring economic policy uncertainty. I choose this example because it's one of the [01:53:38] most influential um papers on text analysis in economics, one of the most cited papers in economics over the past decade. Um and so their measure of economic policy uncertainty computes the [01:53:50] share of articles in a set of newspapers um that also satisfy a keyword query um that indicates that this newspaper article is about economic policy uncertainty. And they perform an audit study that collects extensive ground [01:54:04] truth on mentions of economic uncertainty in a randomly selected set of newspaper articles, right? And so that's another reason we chose this. Um very few papers have really extensive validation data, this paper does. Um and [01:54:18] so what I'm going to show you is the original and debiased EPU indices, where the debiased indices are adjusting kind of the predictions from the text analysis [01:54:30] um towards um what we see in their validation data. Um and I'm going to do that both for their original keyword classifier, and I'm going to apply AI to this problem trained on their labeled data. Um and so you can see what this [01:54:45] looks like kind of with modern text analysis. And so we're going to take 25% of their validation sample and use that for debiasing. So this is the sample where we have both the predictions either of the keyword classifier or the AI, and then we're going to use the [01:54:59] remainder of their data to to train a large language model um to be able to make these predictions. [01:55:05] Okay, so they're taking um tens of thousands of newspaper articles and for each year in their sample, they're constructing the share of newspaper articles that discuss economic policy uncertainty. Um and so you can [01:55:20] see, you know, um their original measure is the pink one. Um they're using keywords. Um we can use a a neural network um and you see it looks not identical, but a kind of fairly similar. Um [01:55:33] and then we have the debiased version of both the neural network, which is the one labeled longformer, the green, um and the keywords. Um and there's a few things to highlight here. Um so the debiased estimates have larger confidence intervals, right? Because [01:55:47] they're taking into account uh the um uncertainty that comes from the mistakes and the text analysis kind of relative to the ground truth uh created by the authors. And the worse the model is, the wider those confidence intervals are. So [01:56:01] the um the keywords aren't as good as the large language model. And so once you adjust for the uncertainty that comes from the prediction error, those are the widest confidence intervals. You know, if we had an even larger set of training data, I think we could kind of [01:56:15] shrink um the um the confidence intervals in the green estimates from the large language model down further. [01:56:22] In the extreme, if you had kind of a perfectly predicted model, um those confidence intervals would look just like the ones where you ignore um that these are generated kind of by a language model and just take it as the truth, which is what the blue and the pink do. [01:56:36] Okay, and you see a little bit of downward bias kind of relative to their ground truth, but it's not huge. And we wouldn't expect it to be huge given how much emphasis this paper put on validation um and how widely cited it is. [01:56:50] All right, so that's a motivating example. I'm going to show you how those intervals are constructed. Um I do want to acknowledge that there's is a very large literature on debiasing. [01:56:59] Um I'm not going to have time to explain what all these different frameworks do. [01:57:02] Some of these are from statistics. Um and essentially here we have different things that your framework may able be may be able to do. What type of estimates is it? What What are the assumptions? Does it allow for an [01:57:15] unknown um pi, which is the function with which data is annotated? So if you didn't know how the data was annotated, you might not know what that is. You might have to estimate it. Does the framework allow for that? Does it allow annotation to be non-uniform? So you don't just have a simple random sample, [01:57:30] which is important. Often a simple random san- sample will not be sufficient. Um does it worry about efficiency? Does it allow for fine-tuning, sample splitting? Does it allow for high-dimensional features, so unstructured data at all? And importantly, does it allow for [01:57:44] aggregation, um which is a point I'll return to in a few minutes and is one of the things that motivated me um to kind of to to work on this framework. [01:57:53] Essentially, at least in the sort of economics I consume, I'd say 99% of the time um you're taking predictions from AI models and you're aggregating them. You know, the economic policy uncertainty predictions is at the level of the [01:58:07] newspaper article and you're aggregating that up to economic policy uncertainty in the US in 1997 and then taking a log and interacting it with firm-level exposure. You're doing all kinds of things to it. And these frameworks, kind of existing frameworks, don't allow for [01:58:21] that. They require you to kind of have validation data at the level of your parameter of interest. And so essentially the summary here is lots of these pieces existed before, um but we're putting them together um into one framework so that you don't have to go [01:58:35] read a bunch of different papers. It's all in one place. [01:58:38] All right. And so I'm going to introduce that framework now. [01:58:42] Um and so we're going to start with uh structured data, which we refer to as M. [01:58:50] These are low-dimensional data that can be used directly in estimating equations, but we don't have them initially. All we have is the unstructured data, the text, the images. [01:58:58] Uh they're high dimensional, we can't use them in our analysis. And the those structured data are missing, and so we need to impute them. [01:59:06] And so we're going to impute them with some function um mu hat. Um because, you know, we have this criterion-based measurement that we can use for annotation. Like that could be me sitting there looking at the newspaper article and saying, "Does it, [01:59:20] you know, given this criteria, does it match? Is this about economic policy uncertainty?" You know, I can't label 100 million newspaper articles, um but a neural network can do that pretty easily. Um and so we're going to have the scalable technology mu hat [01:59:34] um uh that allows the researcher to leverage the full kind of unstructured data set, which might be on a massive scale, um but, you know, it's a black box. Um and deep neural networks are [01:59:48] increasingly serving as the imputation function. [01:59:51] Okay, so we're going to use potential outcomes notation, not surprisingly, as I said, this missing data problem is closely related to the Rubin causal model, which is itself a missing data problem. Um and so we have our structured data M, [02:00:06] um and we have, you know, the ground truth that's just what comes from applying your rubric, that's M star, and we we observe that if A, which is the annotation indicator, equals 1. [02:00:17] And so we're going to need uh some assumptions for this framework. You're going to see that they're very similar to the causal inference assumptions because again, they're both missing data problems. So we need consistency of potential outcomes. All that means is that annotation status is well defined, [02:00:31] which will tend to hold trivially, and that the label for any given kind of piece of any given instance, so any given text, image, etc., depends only on its own annotation status, not on the annotation status of other uh observations. [02:00:45] Um you can do that by having a well-defined rubric, right? That should help with that um that assumption. The key assumption is missing at random after adjusting for observables, which we denote here by X, the annotated data and the unannotated data are comparable [02:00:59] in their ground truth values and the values that we you would get by applying the rubric. And so this is just analogous to selection on observables in causal inference. Um, it's the um it's essentially the same assumption. You could call this annotation on [02:01:12] observables. The third assumption is that we have a known and bounded annotation score function. This annotation score function pi is our decision rule for deciding which instances we're going to label. [02:01:26] Um, this embeds the assumption of strong overlap often seen in observational causal inference setting. So we're not putting zero probability on annotating certain types of uh text or images. Um, we're going to assume in the baseline [02:01:41] it's known um and it will be if you're the one who annotated the data, you know you know how you did it. Okay, so this is pretty weak um but we can relax it um for settings where it's not known. Um, I do want to make a note um that some people say like I don't like this [02:01:56] assumption of strong overlap in this setting because let's say you have a billion texts. I mean realistically you can never annotate more than a vanishingly small fraction of them if you have truly big data. Um, and one of [02:02:09] the kind of advantages of seeing that this is analogous to the causal inference setting is that there's a causal inference literature on decaying overlap and from that we can see that nothing would kind of fundamentally change um if we allowed the number of [02:02:23] annotations to be diverging but the ratio of labeled data to unlabeled data to converge to zero, right? So even if you're in the world um where uh you have a billion things and you can never label more than a small share of them, you [02:02:37] know, you should be okay. The size of the labeled data is going to be important. Um, I'll just make that as an aside. [02:02:45] The final assumption, which we only need if we care about efficiency. We don't need this to be unbiased. [02:02:51] Um, is that we want the expected square error of our estimator to go to zero as the amount of data we train the estimator with goes to infinity. And this is very weak in the context of neural networks like transformers. Um, [02:03:04] but again, um, you don't need this assumption for efficiency, right? You could use a model off the shelf. It's not doing anything asymptotically because it's fixed and you're fine. Estimates will still be unbiased even if the efficient thing to do would be to train a model to get it [02:03:19] to be as accurate as possible. Now, if the annotation score function is estimated, so you're not the one who labeled the data, you have to estimate how somebody else did it, um, then you need to make assumptions about the rate of convergence. But we're not going to [02:03:32] worry about that, um, for now. And I think they're reasonable. [02:03:36] Okay, I want to spend, uh, kind of the remaining time discussing common empirical scenarios. Um, and again, everything that I'm going to show you, um, there is an R package for which is linked at the end of the slides. [02:03:49] Okay, so first of all, I want to talk about, how do we identify a mean with missing data? Um, you know, so we want to we've, um, ran the language model over the newspaper articles and we want to know mean economic policy uncertainty in each [02:04:02] year, like in the figure I showed you. So in all of this, we're going to follow, um, Chen et al. 2008, um, who provides general results on semi-parametric efficient estimation for parameters identified by moment conditions with missing data. And so we [02:04:16] can put everything into her framework and it kind of makes things the most straightforward. Um, and so we can derive an expression for that mean, um, which some of you will recognize as an AIPW estimator. Um, for others, this [02:04:30] might look, you know, just, like, um, a bunch of confusing math. And so I want to go to the next slide and break this down a little bit. So, remember we're trying to estimate a mean from predictions from an AI model. [02:04:44] Um and so, we can break this expression down into two terms. [02:04:49] Um the first term A is just the mean um from your AI predictions. This is what you would get if you just ignored the fact that you use AI to generate this data, you know? So, this is like kind of the blue and the pink estimates in the [02:05:02] figure of economic policy uncertainty that I showed you. This second term B is essentially an estimate of the measurement error of the neural network in the annotated sample. [02:05:14] Um and so, you know, if uh your if your LLM makes perfect predictions, this term will be zero and it will go away. On the other hand, um if your um LLM makes very poor predictions, this term is going to be very large, it's going to have a [02:05:29] large variance, and it's going to blow up your standard errors. But, your estimate is going to be adjusted kind of back towards your annotated data, kind of waiting according to the probability um that that type of observation with those observables appears in the annotated data. [02:05:44] We could write this a different way where instead we think of it um as uh the mean in the annotated data plus a non-parametric regression adjustment to adjust for systematic differences between our large unlabeled sample. You [02:05:58] know, so the first way of writing it appears in a large literature on black box AI. This way of writing it appears in a list at all paper. But, they're the they're two different ways of thinking about the same thing. [02:06:08] Okay, if you understand mean estimation, you will understand every other estimator that I'm going to quickly show you. Essentially, in all of these, in our annotation sample, we're measuring the measurement error, and then we're adjusting our estimates accordingly, and [02:06:23] that intuition carries through to many different examples. [02:06:26] So, we could do linear regression um and think about the, you know, the standard expression for an OLS coefficient, right? R prime R inverse R prime Y. [02:06:35] That's what we have here except the Y, which is our missing outcome, now we're doing the same adjustment that we did for the mean, right? Where we're adjusting um for the fact that our annotation data kind of differs systematically from the [02:06:50] predictions um from our AI model. And so, you can think of this just as the standard OLS estimator, but the outcome has been replaced with what's sometimes called a pseudo outcome um in this literature. Um one important thing to note here is that [02:07:04] your imputation function um should be a function of your context-specific variables. So, like the controls in your OLS regression, if they are relevant um for uh for predicting the outcome, which is something I think that hasn't been emphasized in the literature. [02:07:18] Okay, we can do linear regression with a missing regressor. This looks much messier, but it's still like the standard OLS regression coefficient um estimator, but now um adjusting for the fact that our regressor is missing. [02:07:33] We can do IV. Let's say our outcome is missing. This is our standard TL TSLS estimator, but we're adjusting the outcome for the systematic differences between the imputation and our annotated data. [02:07:46] Again, we can make the instrument missing. Um looks a bit messier, but same thing. Um standard TSLS, but adjusting for the missing outcome. We can do differences in conditional means. [02:07:57] This is where differences in differences comes in. Regression discontinuity, average treatment effects in RCTs. And essentially, these are all just differences in conditional means. And so, the estimator is just going to be a difference that kind of in the conditional mean, which is computed [02:08:11] analogously to the mean example that I showed you. Um we can do the same thing with a conditioning variable. Again, it looks messier, but it's this you know, the standard estimator, but adjusting for these differences. And finally, we can [02:08:24] do this much more generally. Um I know that this kind of looks much messier with the notation. Um, and so I don't want you to worry too much about the expression, although this expression there, like it should look familiar, right? It's just a more general form of [02:08:39] the expression we've seen for every one of these estimators in the just identified case. You can actually do the same thing in the over identified case, although I won't talk about that today. [02:08:47] Um, there's two conclusions that come out of this that are important. So, first of all, this has a type of double robustness property. What I mean by double robustness is we can use kind of any imputation function we want. You know, you could you could apply AI in [02:09:01] the stupidest way possible, be super noisy, you know, your confidence intervals would be massive, um, but they would still be valid when you do the debiasing. Um, and the second thing I want to emphasize is that the efficiency implications of this are you want it [02:09:16] seems obvious, but, um, you know, this is what comes out of the proofs as well as that you want to learn that imputation function as well as you can. [02:09:25] Um, you want your neural network to be as accurate as possible, and that will give you kind of the most efficient estimates. [02:09:34] All right. Aggregation. [02:09:36] Um, as I emphasized a few minutes ago, the bedrock assumption in this literature is that ground truth data are available for the imputed variables that are used in the estimating equation. [02:09:47] This typically fails cuz we might take, you know, say millions of individual social media posts or newspaper articles, aggregate them up to some aggregate level, we might transform them, like frequently we take a log, we might interact them with something, and then we run our regression, right? And [02:10:02] that's really kind of what motivated me to think about this, um, because it's a mismatch between kind of the existing work and what often applied economists need to do. Um, and fortunately, the answer is really easy. Um, so this is [02:10:15] this is great news for all of us. Um, and so, um, we're going to kind of assume you want to run some regression, right? And it's on some kind of aggregation of your data that you've [02:10:30] imputed with AI. Um, and so think of this as mean economic policy uncertainty kind of transformed in some way. Well, what you can do is you estimate the aggregation um, using the Mars first step estimator. [02:10:44] So in this case using the mean estimator that I showed you. So we estimate mean economic policy uncertainty using that mean estimator. I showed you those estimates in the graph. Um, and then we can plug that into our aggregated [02:10:59] regression. Now, when we debias does that mean we don't have measurement error anymore? No, it doesn't mean that. [02:11:04] We still have measurement error, but what valid debiasing does is it takes um, systematic measurement error and turns it into classical measurement error. And so our estimate of mean economic policy uncertainty in the US in [02:11:17] 1997 still has measurement error, but because we've used um, the Mars framework to debias that, now that measurement error is classical. And then we can just proceed um, kind of with um, the standard adjustments for classical [02:11:32] measurement error. You can do that in Stata. People have been doing this for decades. It's all very straightforward. [02:11:37] Some times people don't like adjusting for classical measurement error because they say the variance of the measurement error, do we really know that? Well, the situation where we it's plausible that we know that is that when we have a really large sample and in the big data world that's precisely where we are. And [02:11:52] so I think it's kind of plausible to do this adjustment. Um, and it's because it's classical measurement error, we know how to address it. So we can extend it to clustering, to panel data, um, to heterogeneity in the variance, [02:12:05] um, to kind of all types of different things, right? And so it makes the problem relatively straightforward, but what is essential is that we're going to when we compute that aggregate measure, we're going to debias. And at that [02:12:18] level, we do have the validation data to create that mean or whatever function kind of we're taking of the kind of individual level predictions to aggregate up, you know, those newspaper article predictions into [02:12:30] something that is for the US in 2002. And so we can do this for Baker, Bloom, and Davis. They're going to look at how the change in employment relates to the change in log economic [02:12:42] policy uncertainty interacted with firm year policy exposure plus some additional controls, etc. [02:12:51] And this is what we get. So, on the right-hand side, the blue and the green are kind of the estimates ignoring that we have imputed anything, you know, with the keywords and with um [02:13:05] with the LLM. If we take, you know, on the other hand, we debias to compute mean economic policy uncertainty, but we ignore the classical measurement error problems, we get the gray and the pink that are on [02:13:19] the right. So, that's again with the keywords and with the LLM. And then finally, we can adjust for classical measurement error on the left. It's blowing up the standard errors. Part of this is because they're taking the difference in economic policy uncertainty, which makes the classical measurement error [02:13:34] correlated across periods. But in general, in this case, the conclusions are actually pretty similar. If we had a slightly more accurate neural network, which I think we could with more training data, you know, we would get the same kind of statistically significant results. But this is not the [02:13:49] paper I'm worried about that I reproduced cuz they had really, really extensive validation data and were really careful to try to measure things accurately according to their construct. [02:13:57] You know, I'm I'm more concerned in cases where people aren't doing that. Um that we need a form of validation. [02:14:05] Okay. You know, so I'm running short on time. [02:14:09] I'll just say that, you know, in the paper, we talk about how do you annotate data in practice. Okay, but I'm going to kind of skip through this um because I'm short on time today. You can do a power analysis. How many labels do you need? [02:14:22] The important thing is that in practice we find often times you don't need that many if what you're trying to predict is a binary indicator. You know, often times a few hundred labels is enough. [02:14:32] Okay, so this is not something that's prohibitive in terms of the amount of time it takes you. Hopefully all the time that you're able to save by maybe having coding agents help you run your regressions should be plenty of time to be able to do some validation. [02:14:45] Um All right. [02:14:48] Having more annotated data is good kind of but the most important thing is having a more accurate model uh for your standard errors. [02:14:55] Um so in conclusion, deep learning provides powerful tools for measurement um accounting for imputation bias avoids many challenges that arise with imputing when we just treat things as proxies. Um And we have a package to hopefully make [02:15:10] this easy for everybody to implement. Um but of course um this framework does not apply um to everything. It applies to many common empirical scenarios but notably what happens if the missing at random [02:15:23] assumption um is impossible to satisfy um and that's what the final lecture is going to treat. Thank you. [02:15:30] >> [applause] >> Okay, awesome. Um we're almost there. It's like after 5:00.