Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # Measurement in the Age of AI: Discovery, Definition, and Observation, Melissa Dell, Harvard University and NBER Authors: Discussant: None Video: https://www.youtube.com/watch?v=ydYXe2FxErk&t=0s ## Talk (00:00:00 – 00:41:42) [00:00:02] Thanks to the organizing committee for inviting us and so we've divided this lecture into four parts and in the first part of it I'm going to focus on measurement in the age of AI discovery definition and [00:00:17] observation. And so our starting point for this entire lecture series is that AI is transforming a lot of things but in particular for our purposes today it's [00:00:30] transforming measurement. So AI systems can convert high dimensional unstructured data by which I mean data like text or images audio video into low dimensional variables that we can interpret and that we can use in [00:00:44] statistical analysis. We wouldn't plug a book directly into our regression but we might measure something about that book you know with the with the language model and then be able to analyze that. Um This can often be done cheaply and at a [00:00:59] massive scale. Um and so understanding the implications of AI for empirical analysis requires first examining how AI reshapes the measurement process. [00:01:11] And I'm going to start by asking maybe what seems like a silly very simple question which is just what is measurement um because I found um in working on these topics and discussing them with people that often [00:01:25] times are kind of differences in understanding and in evaluating the challenges raised by AI is a difference in understanding what measurement objectives are. And so I want to start by answering this kind of very basic [00:01:40] question in the context of social science what is measurement? And so the core idea is that measurement is essentially representation. It maps a complex high dimensional reality into [00:01:52] lower dimensional variables lower dimensional measures that we can then understand and analyze. And so it's inherently going to preserve some features of reality, but discard or transform many other features. [00:02:07] And so often times in our daily life, we might think of measurement as kind of a passive observation, you know, how many degrees are on the thermometer, how many cups are in the measuring cup. But when it comes to social science research, [00:02:20] measurement is really not passive. It's a disciplined simplification of reality. [00:02:26] Um and so why is simplification necessary? Well, reality is vastly high dimensional. We see this across, you know, literary, philosophical, religious traditions. You know, so Shakespeare said there are more things in heaven and earth than are dreamt of in your [00:02:41] philosophy. From Borges, we have to think is to forget differences, generalize, make abstractions. And measurement formalizes this simplification. It specifies what features of reality are preserved in the [00:02:56] variables that we analyze. Um and so some of you may have heard the expression the map is not the territory. [00:03:03] In particular, Borges imagines cartographers who pursue perfect representation until they produce a map whose size was that of the empire. But this perfect map is actually useless [00:03:16] because it no longer reduces complexity. It ceases to function as a map. [00:03:23] And so this selectivity, it doesn't make measurement merely subjective, but rather the relevant question is does a given simplification preserve the features needed to support useful inference about the underlying high dimensional world for whatever your [00:03:38] research question is. So we think of measurement in short as a disciplined, fallible, and purpose-oriented attempt to learn from reality through simplified representations. [00:03:49] Mathematically, we can think of measurement as a projection that maps this high dimensional reality into lower dimensional representations that we're going to interpret and that we're going to analyze. Um, and because X is high [00:04:02] dimensional and inherently there are many many ways to map a given concept to a low dimensional measure and naturally these may preserve different features of reality and thus they may support [00:04:15] different empirical conclusions. And so in particular, I find it helpful to think about the different stages of measurement and these lectures are going to be organized around the different stages of [00:04:28] measurement and the implications that using AI for each stage has for downstream estimation. [00:04:35] >> [snorts] >> And so first of all, we have discovery. [00:04:37] Given the complexity of the world, what features exist and are worth measuring? [00:04:44] Then we have construct definition. How should the relevant concept be operationally defined? [00:04:51] And then finally we have observation. How can that definition be implemented at scale? And each stage of this measurement process involves different types of uncertainty and they have traditionally been examined through different disciplinary and [00:05:04] methodological lenses. Um, so let me elaborate a little bit. [00:05:09] >> [snorts] >> Okay, so stage one discovery involves uncertainty around which patterns, dimensions, or relationships exist and are worth measuring. Um, methodologically we would typically pursue discovery through theory or [00:05:22] contextual knowledge. We might do exploratory analyses, but increasingly unsupervised learning and data driven discovery um, are being used um, to um, to to drive discovery. [00:05:37] Okay, in stage two, construct definition, the uncertainty is about how a phenomenon that's of interest can be translated into some operational construct for a given research setting. [00:05:49] So, maybe you're interested in studying economic policy uncertainty and you have newspaper articles, how do you actually operationalize that so that you can measure economic policy uncertainty from the news articles? And again, oftentimes [00:06:03] we um come from theory, we come from domain expertise, maybe from piloting and even qualitative validation. Um oftentimes we make do with whatever externally available proxies are available. Um you know, maybe [00:06:17] expropriation risk for foreigners wouldn't be the perfect way to measure property rights, but it was the available way and so kind of we we do as best as we can with with what exists. Um the formal literatures on construct [00:06:31] definition are mostly in psychometrics um and to a less a lesser extent survey design. [00:06:37] And then finally, there's observation. How accurately can the construct be observed at scale and how do errors propagate to the target parameters? [00:06:46] Um and so, this is where statistics econometrics is quite well developed and we look at this question through the lens of measurement error and semi-parametric inference with missing data. [00:07:00] Okay. A final point that I want to emphasize about measurement is that traditionally we've treated it as fixed typically. Um measurement costs oftentimes prevent researchers from comparing different measurement functions. Um and therefore, [00:07:15] we've often relied on externally constructed measures as fixed proxies for things that are theoretically meaningful. And so, naturally the focus is on justifying the proxy um rather than on comparing alternative measurement functions. [00:07:28] However, AI as a measurement tool um is is changing this and quite rapidly. And so, we can think of deep neural networks as flexible implementations of measurement functions uh that learn structured representations from [00:07:43] high-dimensional inputs such as text, image, audios, and video. Um, and these representations can then be interpreted and used in statistical analysis. And this is changing the measurement landscape by making it much more [00:07:56] feasible, um, to use unstructured data, um, to construct variables for our analysis. So, we can measure things that we could have never measured before, or maybe we could have measured them, but now we can measure them on a much larger scale. [00:08:09] >> [snorts] >> And so, before neural networks, researchers often extracted low-dimensional features from unstructured data using hand-engineered rules. Um, neural networks take a different approach. Instead, they learn representations from empirical examples. [00:08:23] And so, essentially, they use millions to billions of learned parameters to transform the inputs, which are things like text or pixel data, which are inherently high-dimensional, into lower-dimensional representations. [00:08:35] Examples would be, you know, you might transform it into a continuous vector, or you might transform it into a predicted label. I have a newspaper article, does this newspaper article mention economic policy uncertainty, yes or no? And so, the input is the article [00:08:49] text, and the, um, output is that yes, no label. That's a low-dimensional representation that's easy to include in regression analysis. [00:08:59] And so, neural networks are trained to perform specific tasks. They predict the next word in a sequence in the case of a language model, they might classify an image in the case of a vision model, identify objects, match text to labels, [00:09:12] et cetera. And their parameters are adjusted through gradient descent. And so, the key idea behind neural networks is that if you have sufficient training data and an appropriate objective, you can learn which features of the high-dimensional data that you need to [00:09:27] preserve to perform that given task. Why are neural networks so powerful as measurement tools? Well, first of all, they generalize beyond the task which they were trained on. Language models are very useful for predicting the next [00:09:41] word, but they can also perform a variety of other tasks. Um, you might need to tune them on some additional data, but much much less data than you would need if you were training it from scratch. And so you can take a general purpose language model, train to predict the next word, and use it for all kinds [00:09:55] of different tasks. Um, across many tasks, neural networks outperform human engineered rules. Um, and then finally, neural networks are highly scalable. And so that means that we can substantially lower the cost of converting [00:10:09] unstructured information into low-dimensional variables. [00:10:15] I do want to caveat here. I don't want to sound overly optimistic. Um, in many scenarios, credible AI measurement still requires substantial fixed investments. [00:10:26] This is true for a lot of the work that I do with historical documents. If you want to achieve sufficiently accurate predictions, you may need to develop high-quality training data, and then fine-tune tune a customized model. And [00:10:39] so I don't want to imply that researchers, as a requirement, need to be comparing many different potential measures if it remains very costly to use AI to construct the measurement function, you know, that can then be [00:10:52] applied at scale to a particular measure. However, there are other settings where off-the-shelf or very lightly adapted models perform quite well. And this makes it inexpensive to construct many candidate measures of the [00:11:07] same broad concept. And so in this setting, the bottleneck shifts from developing any type of scalable measure to choosing amongst many plausible measures. And these measures may yield different estimates of the target parameter. [00:11:21] And so this, in my view, closely parallels the empirical revolution that came with dramatic declines in the cost of personal computing during the 1990s. [00:11:32] And so, cheap computation enabled researchers to estimate not just one empirical specification, but hundreds, thousands, or even millions. You know, the I just ran 2 million regressions problem. Um and this forced the field to [00:11:45] ask, what should discipline the choice amongst possible specifications? And what makes a specification credible? [00:11:54] And in response to cheap computations, large literatures developed around topics like causal inference, robustness, multiple testing, et cetera. [00:12:02] Um and so, AI-enabled measurement raises an analogous problem. By expanding the set of feasible measurement function, it underscores the importance of assessing what makes a measurement credible, and how do we choose amongst multiple [00:12:17] measurements? When measurement is costly, you know, we might just inherit a proxy, adopt an existing definition, implement the only scalable procedure available. With AI, we can expand measurement possibilities [00:12:32] at each stage of the pipeline. So, in discovery, we can identify patterns or dimensions that we the researcher doesn't have to specify in advance. With construct definition, we can generate and compare alternative operationalizations. [00:12:45] And with observations, we can implement a chosen definition at scale. [00:12:49] And so, I want to elaborate um on how AI is influencing each of the stages of measurement. Uh discovery, traditionally guided by theory, contextual knowledge, incremental advances in the existing literature. And this works quite well [00:13:04] when the problem is sufficiently understood for pre-specified measurement to be met well motivated. Uh but the risk is if these are the only tools that we have available, um we risk looking for the keys under the lamp post. That is to say, we only search for concepts [00:13:18] that we already know how to specify. With data-driven discovery, the researcher specifies a representational procedure rather than a predefined construct. And so they don't say, "I want to study economic policy [00:13:33] uncertainty because this theory says it's important." Instead, they start with a procedure that allows them to discover, you know, what are these what topics do these open-ended survey questions focus on? And so examples of methods might include embeddings, [00:13:47] clustering algorithms, dictionary learning methods, autoencoders. I know I haven't defined any of this and Ashish will be talking more about um discovery um after this lecture. [00:13:58] Um But in short, these procedures summarize high-dimensional data by identifying structure that you the researcher do not have to specify in advance. [00:14:09] All right. So these discovery procedures might identify latent patterns, clusters, dimensions, categories, recurring components. Um some methods group similar observations um together and others learn continuous dimensions [00:14:22] along which observations vary. Um This discovered structure can enter measurement in two ways. Um so oftentimes this discovered structure is informative, but sometimes it might be kind of a bit noisy or loose. You might [00:14:37] want to use it as an input into a more precise construct definition. Um and you're going to use kind of the output from a discovery process uh to formulate a more precise concept. Importantly, you'll need to measure that on a [00:14:50] separate sample to reduce post-selection inference concerns. [00:14:53] Um or maybe this discovered structure is itself the measurement output, in which case you're going to interpret it through interpretability methods, high-dimensional statistical analysis, and perhaps external validation. [00:15:07] So discovery is useful when as a researcher you're unsure what to look for, you suspect existing conceptual categories are incomplete, or as a way to discipline discovery by pre-specifying a data-driven procedure [00:15:21] uh rather than pre-specifying what you're going to measure. And so, I think a great example of when this might be useful is let's say you're going to run an experiment and collect a survey afterwards. You want to have an open-ended uh response question. If you [00:15:35] pre-specify what you're looking for, it kind of defies the point of having an open-ended response. Uh but what you could do instead is pre-register a procedure to identify recurring themes in the open-ended survey responses. And that [00:15:49] way you keep the spirit of an open-ended response, but you can still pre-register. [00:15:55] All right. So, construct definition. How should a given concept be operationalized for a particular research setting? [00:16:02] And this is actually quite difficult um if the concept's multi-dimensional, not directly observed, closely related to other phenomena. So, for example, we might want to measure trust, but what is that? It's some high-dimensional latent [00:16:15] concept. Um you want to capture whatever dimension of trust is relevant to the research question and not simply proxy things like institutional quality, income, or other correlated variables. [00:16:29] Okay, so AI is not going to determine construct definition on its own, um but it can reduce the cost of data construction um if zero-shot models um perform well. It can both expand the [00:16:42] candidate set of measures that researchers can examine and the empirical evidence um that we can use to choose choose amongst them. [00:16:51] Okay, so the AI needs to be combined with theory or with domain expertise that clarifies what the construct should capture for the research question, um what it should be distinct from, um [00:17:02] where it should be stable, etc. So, in observation, the role of AI is to serve as a scalable measurement technology. [00:17:13] In observation, the researcher has a specified kind of definition in hand and they want to apply it repeatedly to high-dimensional inputs at a low marginal cost and AI is ideal for this. [00:17:27] It makes it possible to classify large text corpora to extract information from vast collections of images, satellite data and more generally to construct low-dimensional features from [00:17:38] unstructured data on a vast scale. Um And so importantly though, there's choices that you have to make. You know, choosing the model, the prompt, the training data, input [00:17:52] distributions apart from defining the concept itself, kind of taking the construct that's given, you still have to make choices to implement the AI and those can produce different measurements. [00:18:05] If there are errors, they are unlikely to be classical. You get systematic biases from the network architecture, from the distribution of training data, from implementation details. A neural network is taking non-linear transformations at every level of the [00:18:19] neural network. That violates classical measurement error. People are often times doing binary or multi-class classification. That violates classical measurement error. There's no reason to expect this measurement error to be classical and if the AI is making biased [00:18:34] predictions, that will propagate to estimates of the target parameter. So an implication of this that we argue is that validation is central to AI-enabled measurement. [00:18:45] Okay, so I want to turn now to first saying a word about the role of validation more generally in measurement and then talking about validation across the measurement pipeline. [00:18:57] Okay, measurement remember does not aim to reproduce reality. So validation cannot show that a measure perfectly captures a latent concept or recover some fundamental truth. You know, sometimes I've had people push back [00:19:11] against the importance of validation because they say, "Well, that's a latent concept. You can't show that you've accessed the truth, so why are you doing this?" Well, we have kind of more modest goals that are still useful. [00:19:24] Um And so, measurement is aiming to construct a simplified representation. [00:19:30] In principle, we would like to validate whether the measurement is useful for inference about the high-dimensional world for economic questions. That's often quite difficult, at least in the short run. The [00:19:43] underlying reality is complex. It's partially unobservable. It's multi-causal. It's not easily manipulated. And so again, we have kind of a more modest goal. Um We want to [00:19:56] judge whether measurement provides a useful guide to reality. But to do that, what we're going to do is state an explicit measurement criterion and assess how well the measure behaves [00:20:09] relative to it. And this is true kind of at all stages of measurement. And you know, this criterion is stated explicitly. It can be interpreted. It can be audited. It can be debated. It can be refined. [00:20:24] Okay. And so, some people again might say, "Well, why what's the value of anchoring to explicit criteria?" I think that this is particularly valuable in the age of AI because it addresses some of the key [00:20:37] limitations of AI. And so, with neural networks, the internal mappings from the input, so let's say the text that you put into the language model and the outputs, you know, the prediction of what's the topic of this article, are too complex to [00:20:52] understand by inspecting the parameters. You know, if it's an open-source model, you could open it up and look at the millions to billions of parameters inside, but that won't help you to make sense of it very easily. Learn criteria are not expressed as explicit audible [00:21:06] rules. Um, and the biases of neural networks shift in complex ways with the inputs. Uh, humans are not good at predicting these shifts, at least compared to how good they are at predicting how the accuracy of human shifts uh with the input distribution. [00:21:21] And so, validation is making AI interpretable not by elucidating its inner workings. You know, if we as economists could do that, we could answer all the biggest questions in AI. [00:21:31] That's a very challenging problem. Instead, what we want to do is make AI interpretable by evaluating its outputs against criteria that are specified by the researcher um and justified by the researcher that are important uh to [00:21:46] answering the research question at hand. Um And so, validation also addresses the abundance of measures that AI can cheaply produce. [00:21:56] Um Measurement choices, when you can produce many measurements cheaply, um become objects of empirical empirical comparison. And so, validation provides explicit evidence for choosing amongst candidate measures, for correcting or [00:22:11] refining measures, um and for deciding when additional investments like maybe training your own customized model, training a larger model for longer or more data. These are all costly investments that will improve neural network predictions. Um, validation can [00:22:26] help you decide if these investments are worth it or not given your measurement objectives. [00:22:33] Some people will criticize validation cuz they'll say, "Whoa, wait a minute. [00:22:36] Why um why are you privileging humans when humans are noisy, error-prone, etc.?" Well, validation is not privileging humans per se. Um, it's privileging explicit interpretable [00:22:49] criteria over implicit decision rules learned by a black box model. So, yes, the validation data can come from human-coded labels. They might also come from scientific instruments, from administrative data, from theoretical predictions, from algorithmic criteria, [00:23:03] etc. Even when a neural network is prompted or trained to apply a given criterion, its internal decision rules are not directly interpretable, and there's some evidence that while sometimes it does [00:23:17] this well, you know, oftentimes you give it a specific a set of criterion and it can actually drift back to kind of what it's learned in pre-training, which is not necessarily what you may be after for your research question. [00:23:31] You know, another potential critique of of validation is it's imperfect. There will inevitably be edge cases about which reasonable people may disagree. [00:23:41] Um this, you know, does not imply that validation lacks value, but it identifies where refinements may be needed. Um A neural network will encode richer information than a criterion that you [00:23:55] state, but the question is, does this suggest a better interpretable construct, or does this richer information mostly encode, you know, things that are irrelevant or that are confounding variation in the context of [00:24:08] your research question, or that are difficult to interpret? [00:24:14] All right. Um so, finally, I'd like to speak about validation across the measurement pipeline. [00:24:22] Okay, because the three stages of measurement involve different sources of uncertainty, validation will take different forms at each stage. So, discovery asks, has the exploratory procedure uncovered [00:24:37] meaningful structure? In construct definition, we ask, does the operationalization capture the intended dimension of the concept? And in [snorts] observation, we ask, does the measurement procedure [00:24:50] accurately implement whatever criterion we selected through construct definition? [00:24:57] All right. Um so first I want to talk about validating discovery. Um so in data-driven discovery, the researcher has not yet specified a fixed criterion against which the output can be [00:25:10] validated. The entire point of it is you don't want to say I'm looking for these particular things. Instead, you kind of want to let the data speak about what it contains. [00:25:20] Um and so with validating discovery, we want to ask has an exploratory procedure uncovered meaningful structure? We want it to be meaningful enough that we can actually interpret it. If we can't interpret it, what's the point? We might [00:25:35] want to refine it into a construct, or we might want to use the output of the discovery process as the measurement object itself. [00:25:45] Okay, so the most kind of straightforward way that we would uh validate data-driven discovery um is internally. [00:25:55] We want to know do the data actually exhibit separable structure under a particular algorithm or representation? [00:26:02] Um so how you validate this will depend on what method you're using to discover patterns in the data. For example, if you're using clustering, uh so you're grouping kind of texts into different clusters, um you might say, well, [00:26:16] uh does the output show within cluster cohesion and between cluster separation? [00:26:21] Um internally, we also care about stability. Uh so meaningful structure should recur under perturbations of sampling, pre-processing, implementation, et cetera. [00:26:34] All right. Beyond kind of those internal uh validity concerns, a related concern is does the discovered structure just reflect the method that we used to elicit it rather than a meaningful dimension of reality? [00:26:48] Um and so there's a highly cited paper in psychometrics all the way back from 1959 by Campbell and Fiske and they develop what they call the multi-trait multi-method framework for validation. [00:27:00] So they argue that a measure should converge with measures of related concepts, remain distinct from nearby phenomena that you want it to be distinct from for your research question and be recoverable using alternative [00:27:14] methods which is the multi-method part part of the framework. So what would this look like in AI enabled discovery? [00:27:24] Well, you might ask whether the structure that you've discovered, the way that you've grouped the data or learned about the dimensions of the data can be recovered through different measurement routes. You might vary the the representation model like what neural [00:27:37] network or model are you using to group the data? You might vary the discovery algorithm. You might vary the method that was used to elicit or compile the underlying raw data or you might vary the underlying raw data source. [00:27:52] And the more you see the same patterns kind of arise regardless of the exact method you use kind of the more plausible that you are uncovering a real pattern in reality and not just [00:28:04] uncovering some sort of artifact of the method. [00:28:10] Okay. >> [snorts] >> Um you could have a great discovery process, but if you can't interpret the output, there is no point, right? And so a discovered representation is useful when we're able to interpret it. We see [00:28:24] this in older traditions, but it's important as well in modern AI and representation learning. And so you'll see a lot of people use LLMs to cheaply generate labels or summaries or explanations for the discovered [00:28:37] structure. And these methods can be really powerful, but it is important to caveat that the interpretations themselves uh be validated since the models internal reasoning is not [00:28:49] directly inspectable. Uh when [snorts] possible, external validation kind of can also help with the interpretation. Um and so you might just interpret things by collecting a [00:29:03] small manual audit sample, uh but on top of that, uh many papers um will try to do some type of external validation. And so an interpretation of the discovered structure that you propose becomes more [00:29:16] credible if um it correlates with or predicts or remains distinct from external variables. You know, that that it should relate to or not relate to if the interpretation is correct. [00:29:30] And so this external validation data disciplines the interpretation of a discovered structure. [00:29:39] An important concern when you're doing data-driven discovery is post-selection inference. Um uh data-driven discovery can be vulnerable to overfitting, cherry-picking, ex post narrative constructions, etc. So how do you avoid [00:29:53] this? Um well, the first way is to do sample splitting. So you separate discovery from subsequent measurement. [00:29:59] You use one split of the sample to identify candidate structures, and you use those kind of to understand uh the patterns in the data. Based on that, you refine what you want to measure, and then you implement that measurement in a [00:30:12] separate sample split. Alternatively, maybe the discovered structure is the measurement itself. Um and so in that case, um credibility depends on evidence that the [00:30:25] structure is well-defined, stable, robust, interpretable, externally meaningful, and you can use tools like high-dimensional statistics to do direct testing on the discovered structure. [00:30:39] All right. Um so that's validation of data-driven discovery. Um, I want to turn now to what validation means in the context of construct definition. And so in construct definition, we want a theoretically meaningful [00:30:53] concept uh to be concretely defined and operationalized uh for a specific research setting. Um, and um the seminal work on construct definition comes from psychometrics um [00:31:07] all the way back in the 1950s. Uh there's this work by Cronbach and Meehl, which is one of the most highly cited papers um in psychology and psychometrics, uh which argues that many important concepts to social scientists are latent and high dimensional. So, you [00:31:21] know, so examples would be things like trust, respect for authority, social capital, intelligence. You can come up with some rubric to measure them, but how do you know that you've actually captured that characteristic when it [00:31:34] isn't directly observable or verifiable, when it's very high dimensional and you could be capturing many different dimensions of it, etc. Um, and so their key insight is to argue that when direct observation is impossible, validation [00:31:48] must be indirect. Okay. And so, they argue that a construct is validated not by a single observable variable, but by a broader pattern of relationships that's suggested by theory or by the existing [00:32:02] literature. They call this a nomological network. Um, and uh the idea is that you use theory to say, for instance, you know, what the measure relates to. Um, that's called convergent [00:32:16] validity. What the measure should be distinct from, that's called a discriminant validity. And also things like how stable it should be, etc. And so if we wanted to kind of formalize this with econometric language, the way that I think about it, although that you know, they don't formalize it in this [00:32:31] way, but I think about it as analogous to partial identification under theoretically motivated moment restrictions. So, if you have a, you know, a high-dimensional latent concept not directly observable, no single [00:32:44] restriction establishes that you've captured the intended construct, but additional restrictions can narrow the set of plausible measures. [00:32:52] Um, you know, traditionally economists didn't worry about this um because we emphasize things where direct validation was plausible. You know, we measured prices, schooling, employment, revenue. [00:33:03] Um, but this is changing and in recent years economists have increasingly studied latent high-dimensional concepts where construct validity is much more contested. You know, things like trust or social capital or economic policy uncertainty. [00:33:17] Um, and so AI has the potential to improve construct validity by making empirical assessment of things that are difficult to validate more tractable. If it performs kind of well enough off-the-shelf, we can use it to generate [00:33:30] more candidate measures and to also measure observable implications of the nomological network that can be used to test which measures are plausible. But AI also makes it essential to develop empirically grounded methods for [00:33:43] assessing construct validity. Uh, so here's an example. Um, I took economic policy uncertainty, a concept that's defined by Baker, Bloom, and Davis. You know, I took their rubric as an example and I asked frontier language models, [00:33:57] um, GPT and Opus, um, to give me a definition of economic policy uncertainty. I did this 10 times. They go, they search the literature, they come back with something that sounds plausible, and then I implement that rubric. I mean, you can see that these [00:34:12] are positively correlated, but they're certainly not perfectly correlated and it does, you know, matter potentially for uh the downstream conclusions and whether those are the same as the ones reached by Baker, Bloom, and Davis. And so, if we just have this, which is very [00:34:25] cheap and very easy to do and people are increasingly doing this, we have no idea really if this is slop or if this is capturing meaningful variation in the concept. Um [00:34:38] And so, I think, you know, at its core, this moves the problem of searching across regression specifications, you know, from the 1990s to searching across representations of the underlying reality. If we don't have a principled [00:34:52] and widely accepted method uh for construct validity in cases where what we're measuring is not directly observable, um I think it risks encouraging p-hacking with AI slop on the one hand, while promoting nihilism about what can [00:35:06] actually be learned about a high-dimensional world on the other hand. this is already happening to an extent, and I think that this is going to get worse without taking these issues seriously. [00:35:17] Common critique of construct validity is that it depends on sufficiently developed theory, which we sometimes lack. [00:35:24] Um you know, if we don't have any theory at all, we may be better off just doing discovery, right? If we don't have a way to validate, have I actually captured what I'm trying to capture, you do data-driven discovery, you can learn [00:35:38] some things um about what you're trying to measure, rather than developing a construct that is is under-validated. [00:35:45] Okay, in the final few minutes, I want to talk about observational validity. Um and so, I think it's really important to distinguish construct validity and observational validity. And how I came about this is I have some work on observational validity, and I found kind [00:36:00] of the most common critique from applied researchers was, "Well, construct validity is a big problem, and you're not addressing that, so why should I care about this framework?" Well, actually, they're two separate problems, and they're both important. So, I think [00:36:13] of observational validity, which asks, "Does a procedure correctly implement the operationalization that you've chosen?" as in some ways akin to internal validity. Credibility conditional on a specified target. [00:36:25] That's a construct definition in the case of measurement. It's an estimate in a given setting in the case of causal inference. I think of construct validity as analogous in some ways to external validity, whether the target answers [00:36:37] your broader research question. And so to validate observation, we start with the premise that researchers have access to a costly but credible criterion-based measurement technology for a subset of [00:36:52] observations. Um, you know, this could be a measurement instrument, it could be human etc. It, you know, the criterion may be flawed, uh, but it's interpretable and we're going to take it as fixed. Um And so credible criterion-based [00:37:06] measurement is often costly. Um, so it may be feasible only to apply it to a modest sample, but we'd like to use a much bigger sample, you know, so that we have more power and we can do that with a black box technology like neural [00:37:21] networks. Um, we might use an LLM, a computer vision model for satellite data, etc. Uh, this scalable technology doesn't need to be AI, but in this lecture we'll assume that the scalable technology is AI. And so you In short, you have this accurate criterion-based [00:37:35] measurement. This might be, say, ground station data, but it's very costly. And then you have this very cheap measurement uh, that you can scale to a large sample, but it's a black box. [00:37:46] Um, and so a validation sample contains observations for which both the scalable measure and the criterion-based measure are observed. And under appropriate assumptions, this sample can be used to characterize the measurement error. And if you have an appropriately defined [00:38:00] validation sample, it allows you to estimate the target parameter consistently and efficiently in the much, much larger sample, um, even under non-classical measurement error. And as long as your your cheap measurement technology is good enough, you'll get a [00:38:14] lot more power uh, by using the much larger sample. [00:38:18] Um, one note that I want to make is that observational validity is not a sensitivity check. Sensitivity analysis asks whether conclusions change across different observational procedures, whereas validation links measures to an explicit criterion. It provides a [00:38:33] benchmark for interpreting robustness checks. So, these are both important, but one's not a substitute for another. [00:38:39] Um sometimes people will argue, "Well, I ran this through a bunch of AI models and they all agreed. Isn't that showing that my results are robust and this is good enough?" Um unfortunately, um there may be correlated errors across models. [00:38:52] Uh models, you know, AI models are basically all using the same architecture, a transformer. They use similar training procedures. They're basically all trained on a snapshot of the internet. Um you might use similar prompting choices. So, really the takeaway is that agreement across AI [00:39:06] agent measures may reflect shared biases rather than the measurement quality. [00:39:11] This, you know, obviously means that you can't uh necessarily use one as a valid instrument for another. Kind of more generally, there's an emerging machine learning literature on response homogenization that argues that pre-training makes models broadly uh [00:39:25] competent in similar ways. Um reinforcement learning with human feedback pushes them towards similar preferred answers. Um with output variation driven more by um prompting. And there's some social science evidence as well that by varying the prompt, which I think in many ways [00:39:40] is varying the construct, you can get kind of many different answers to the same questions. Um and so I show this um with a rubric we created, you know, is this historical newspaper article about politics? Um we had um uh newspaper [00:39:53] articles double labeled, created a big sample. Um and so that we have the ground truth for our large sample. And then we ran this through different um frontier models as well as kind of testing our own um fine-tuned model at the bottom. Um and you can see that the [00:40:08] prediction error conditional on our ground truth label is highly correlated across models. [00:40:14] Um prediction error is not random. And if we look at the mean share of articles about politics in our large labeled sample, on the far right, that's the oracle. That's the kind of the the truth from applying our rubric. It might be a [00:40:29] flight flawed rubric, but it's how we define politics. And you can see all the frontier models are overestimating the share of politics in historical newspapers. Why are they doing that? [00:40:40] Well, I think it's this critique that, you know, we give them the explicit criteria, but they fall back on their pre-trained knowledge. And so they've been trained on the internet. They might think the historical news article saying, you know, there's a measles vaccine clinic [00:40:53] downtown today from 1967 is about politics cuz people are often making a political statement about measles today, but that doesn't mean it was about politics under our rubric, right? And so what we'd ideally like is to have a way [00:41:06] kind of to do debiasing where no matter what model we use, we get the same estimate. It may be noisier if the model's not very good, um but when we debias, we'd like to be able to get a similar estimate regardless of our model, and that's what I'll be talking [00:41:20] about after the break. Thank you. >> [applause] >> I'm uh really excited to be here. Um and