Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # Discovery and Hypothesis Generation with Unstructured Data, Ashesh Rambachan, Massachusetts Institute of Technology Authors: Discussant: None Video: https://www.youtube.com/watch?v=ydYXe2FxErk&t=2502s ## Talk (00:41:42 – 01:34:42) [00:41:47] so Melissa, you know, really laid the stage for all of the parts of our plans for this uh this methods lecture. Um so I'm going to spend some time talking about how we can do discovery and hypothesis generation with unstructured data. Um and really my starting point for thinking about this problem is just [00:42:02] a broader observation that has gotten me and so many researchers across many fields very excited over the last year. [00:42:08] Um and that is if you look across fields across science, uh AI and machine learning are increasingly accelerating the pace of scientific discovery. You know, just last year we had AI models for the first time winning a gold medal in the IMO competition. And 1 year [00:42:23] later, we have people posting memes about using AI models to break frontier conjectures in mathematics. In biology, AI is being used to predict novel protein structures and more broadly across the sciences, researchers are using these tools to service patterns [00:42:38] that we as humans may not notice ourselves. [00:42:41] So, in a question I've been, you know, deeply excited about for a long while now is that if this is happening in other sciences, you know, what might this look like in economics? If AI can be used to accelerate discovery elsewhere, how might AI accelerate economics research? [00:42:55] And of course, there's not going to be one single answer to this question. [00:42:59] There's many people doing wonderful work thinking on this thinking about this problem. [00:43:04] One answer might be that AI could accelerate existing workflows that we have. So, for example, in empirical research, there's been a lot of really wonderful work trying to argue that we could use AI tools to automate parts of empirical analysis or policy evaluation [00:43:17] itself. In econometric theory and in economic theory, we could use AI to help us formalize arguments, suggest new results and proofs, just as researchers are doing in mathematics. [00:43:28] But what I want to spend the next 40 minutes or so talking about is that AI gives us the opportunity to formalize and accelerate a new type of problem that we can work on at scale. [00:43:38] And that problem is what a burgeoning literature refers to as the problem of hypothesis generation. [00:43:44] And this opportunity rests on really three observations about how AI and machine learning interacts with economics that a large set of researchers have built up over about about a decade really thinking on the machine learning side. [00:43:58] So, the first observation is that in many of the structured domains we study, machine learning algorithms often out predict our best economic theories out of sample. And this is true in some of the most classic experimental settings where we study models of decision-making [00:44:12] under risk. We study models of strategic interactions with game theory. And in light of this finding, researchers have then been excited to try to answer the question, well, if these algorithms are out-predicting our theories out of sample, what signal have these algorithms uncovered that we as [00:44:26] researchers failed to notice? How could we extract theoretical insight from this black box? [00:44:32] The second observation is that, of course, while we care about doing well in prediction within domain or generation within domain, economics is far more than that. We care about trying to generate theories and ideas that generalize to new settings, that allow [00:44:46] us to evaluate counterfactuals, and more broadly understand mechanisms. And there's been a lot of work pointing out that algorithms on their own may fail on this margin. So, we want something else than just an uninterpretable black box. [00:44:59] And the last observation is one that I think everyone in this room has wrestled with, and that is that in economics, we always are confronted with the lamp post problem. We search where the light is in our structured data. The closed survey items we collect, the typical outcomes [00:45:12] we study in a randomized experiment, the fields of an administrative extract, only capture a limited part of the domain we study. And as a result, you know, for a long time, there have been huge incentives for researchers to try to use unstructured data to record [00:45:26] features that we did not measure before in our structured fields. And you know, this is an extraordinarily incomplete list, but there are wonderful examples of of this across every single subfield within economics. [00:45:39] So, together, these three observations about how AI and machine learning could interact with economics suggests a very exciting opportunity. Could we use AI to analyze unstructured data to try to automatically discover what we should be measuring at scale and how it relates to [00:45:53] the quantities we care about? This is an activity that this literature refers to as hypothesis generation. And then, once we've generated these hypotheses, can we use them and analyze them using our standard empirical and theoretical toolkit? Can they give us the fodder [00:46:07] that we can then ask the typical questions we want to study in empirical economics? [00:46:13] So, as I'm going to argue and point out and hopefully summarize for you over the last really three or four years, we've seen more and more examples of researchers in economics, computer science, statistics, data science, a wonderful interdisciplinary mix of researchers building algorithms to [00:46:27] assist us or automate the process of hypothesis generation from unstructured data. [00:46:32] And I want to anchor that conversation around three empirical examples that I'm going to keep going back to throughout the talk. [00:46:39] So, one important area where we've seen in machine learning interact with economics is in a class of problems that people now refer to as prediction policy problems, which start from the observations that many important and consequential decisions we study hinge [00:46:53] on a prediction. So, for example, in pretrial release, a judge decides whether to detain or release a defendant based on a prediction of whether that defendant would commit pretrial misconduct. And for example, in child protective services, an agency decides whether to open a case and provide [00:47:08] services for a family based on whether a based on a prediction of whether the child suffers from underlying maltreatment. [00:47:15] So, the first example is going to focus on the context of pretrial release. And in pretrial release, though judges see more about a defendant than what is captured in any structured administrative extract, after all, the defendant is in the room with them and they get the opportunity to ask that [00:47:29] defendant many questions, machine learning algorithms trained on structured data to predict misconduct risk systematically outperform judges in making these decisions. And this is a point that's been made going back to some classic work by Kleinberg et al. in the mid-2010s and has been followed up [00:47:44] in many different jurisdictions. So, it would appear that judges are getting these decisions wrong, but what exactly are they getting wrong in making these choices? [00:47:52] So, in a wonderful paper, Ludwig and Mullainathan try to answer this question by turning the camera on the judge and instead of predicting misconduct risk, let's predict detention decisions in Mecklenburg County, North Carolina. [00:48:04] And their starting point is the observation that the single most important predictor for pre-trial detention decision is the mug shots. [00:48:11] Above and beyond structured information about the defendant such as their current charge, their prior record, demographics including age, race and gender that you would think are inferred from the mug shots, known facial features in a large psychology literature and even incentivized guesses [00:48:25] about who is going to get detained from the mug shots. In fact, about 78% of the predicted signal from the mug shot is entirely unexplained by structured characteristics in administrative data or known facial features that you might code from this image. [00:48:38] So, it appears that something unnamed in the mug shot is systematically driving detention decisions. What is it? [00:48:44] The next example I want to talk about goes to a different setting in child protective services. When child protective services an investigator has to decide whether to open a case and provide services to a family based on a prediction of whether a child is suffering from maltreatment. And here [00:48:58] again, we can do a comparison of investigators against the machine learning algorithm, but you find something quite different. Uh in this work I did with Jason Baron, Will Dobbie, Richard Lombardi and Joseph Ryan, what we find is that in Michigan, in fact, a substantial fraction of [00:49:12] investigators, about 36% of them, significantly outperform a machine learning algorithm trained on structured characteristics. You can spend a lot of work trying to unpack that results and what you find is that the gap between investigators and the algorithm is not [00:49:25] explained by anything that you have in your structured data. Something else is separating the very best investigators from the worst ones. So, what is it? [00:49:34] Well, of course, investigators do a lot more than what's captured in the structured data we have. They visit the home, they interview the family, they contact teachers and doctors. None of this is captured in the typical structured data available in administrative systems. But, what we do [00:49:48] have available is that investigators record rich unstructured case notes that document everything they did during their investigation. [00:49:56] So, we could then try to ask the question, can we use these case notes to try to generate hypotheses about what differentiates high versus low-performing investigators? Do they do something or notice something differently in the investigations, and does that help us explain heterogeneity [00:50:10] across decision makers that we otherwise couldn't explain? [00:50:14] And as a last example to give you a flavor this is not just a question you can ask in applied micro, let's think about hypothesis generation in studying household expectations about macroeconomic outcomes. [00:50:24] Of course, how people explain the macroeconomy shapes their expectations, and those expectations in turn drive their behavior. So, in a wonderful recent paper, Andre et al. in the Review of Economic Studies, uh tried to understand what are the inflation [00:50:37] narratives people held during the 2021 to 2022 inflation surge. [00:50:41] The narratives shape people's inflation expectations because it affects how they interpret the news. For example, if they think that inflation is caused by an energy crisis, which led to supply chain issues, that might lead them to interpret future events differently. So, they wanted to start by just asking, [00:50:56] what are the narratives that people actually hold about the causes of inflation, not the ones that we as economists might ex ante write down ourselves. [00:51:04] So, to do that, what the authors did is they elicited a lot of open-ended survey responses. As I briefly hinted at, that's an essential step in this activity. An ex ante menu of these inflation narratives would presuppose the very narratives we were trying to discover, and there's no reason to think [00:51:18] that the ones that we as economists might generate are actually the narratives people hold. But, it led to a very painstaking process where the research team then had to read every single response themselves and hand code them into a directed acyclic graph. [00:51:31] This is only become going to become more popular. Uh collecting open-ended responses is increasingly common in economics research, and will become increasingly cheap as there's wonderful work trying to develop AI interviewers to do open-ended unstructured interviews with subjects. [00:51:45] So, rather than having a human try to look at these narratives and generate hypotheses about them, could we hand the responses to some AI pipeline and let it discover inflation narratives automatically? [00:51:56] So, this is a a nascent and extremely rapidly growing space, as I mentioned, that sits across fields. There is examples in economics, medicine, education, computer science, computational social science. So, this [00:52:10] is going to be an imperfect attempt to try to summarize this literature in about 30 minutes. And it's also going to be particularly dangerous because the frontier of AI is moving extraordinarily rapidly. And what was impractical in this problem may not work today. What was impractical a year ago may not work [00:52:24] today. So, my goal is going to be to try to very briefly describe the structure of hypothesis generation procedures, some of the key choices that are required, and some of the challenges that arise. But, I will of course have to skip details just due to due to time constraints, but the slides are online. [00:52:39] So, what's the setup? We're going to think about this problem in a context where we have access to some data X and an outcome that we're interested in studying Y. And each unit of data X is going to be a piece of unstructured data. Think of it as images, text, [00:52:52] audio, etc. So, in Ludwig and Mullainathan, X is a mugshot and Y is the detention decision. In Baron et al., X is the case note written by an investigator and Y is the case opening decision. And in Andre et al., X is the open-ended response and Y is the [00:53:07] respondent's reported inflation expectations. [00:53:11] So, we observe some unstructured data paired with some outcome Y. And as a starting point in this hypothesis generation exercise, the researcher is going to choose a candidate space of hypotheses, script H, that they want to try to generate into. [00:53:24] At a very high level, each hypothesis is going to name a feature of the unstructured data that we could potentially go out and measure at scale. [00:53:32] Roughly speaking, you could think of that as some function that's going to map this piece of unstructured data to some numeric value. And it's often an indicator for whether it exhibits a particular feature. [00:53:43] So, in these problems, we can describe the candidate space of hypotheses at a very high level. In Ludwig and Mullainathan, it is descriptions of facial features. In Baron et al., it is descriptions of narrative evidence themes expressed in the case notes. In [00:53:56] Andre et al., it is descriptions of causal narratives about inflation. But, the core challenge is that we cannot ex ante specify every hypothesis contained in this space. It's an implicit object. [00:54:07] If we could do that, it would effectively mean we could enumerate everything we could possibly hope to discover from unstructured data. And if you're familiar with the broader literature on automation, this is really a version of Polanyi's we see [00:54:21] them, but we can't ex ante list them in advance. And that's going to be a core challenge in building procedures to do this. [00:54:27] So, at a high level, I think of a hypothesis generation procedure as something G that's going to take in input data, pairs of unstructured data associated with some outcome Y, and then return a collection of hypotheses H hat to the researcher alongside with some [00:54:42] evaluation metrics that assess their quality. [00:54:45] So, in Ludwig and Mullainathan, we're going to take a data set of mugshots and detention decisions and return a collection of descriptions of facial features. In Baron et al., we're going to take a data set of case notes and case opening decisions and automatically return descriptions of narrative [00:54:58] evidence themes in these case notes. Uh and in Andre et al. in a reanalysis, we're going to take these open-ended survey responses and inflation expectations and return descriptions of causal narratives of and about inflation. [00:55:10] So, at this point, everything I've said, this hypothesis generator, could just be a black box. It's anything that maps data to named hypotheses. So, in principle, this could be the researcher. [00:55:21] That's really how we have operated for, uh you know, the last decades, hundreds of years. Tomorrow, perhaps this is just a frontier AI model, make a procedure call to OpenAI, and maybe that's our hypothesis generator in the future. [00:55:34] But, so far how this literature has approached this problem is that many of the hypothesis generation procedures we write down share a common structure to them and they typically organize this problem into four stages. There's going to be a first stage where we try to [00:55:46] score each unit of unstructured data. Then we use our scores to try to contrast units of unstructured data that are high scoring versus low scoring. [00:55:55] Then we try to interpret what separates these high versus low scoring units of unstructured data and then we're going to report some evaluations on a held out test sample. So my goal is to try to walk through each stage at a very high level, discuss how these are implemented [00:56:09] in three examples and some of the key challenges and choices that arise. [00:56:13] And what I want to flag up front is that as I mentioned this is a nascent space. [00:56:17] So each of these stages admits many implementations and often the choice of implementation is going to vary depend on the domain and the specific data modality, whether we're talking about images, text versus waveforms. Um and to this point for folks interested in methodological problems there's been a [00:56:32] lot very little systematic comparison of these choices um and I'm going to come back to that at the end. [00:56:38] But what's going to be important is that at the evaluation stage implementations are going to be judged in largely the same manner. We're going to ask do the generated hypotheses correlate with the outcome of interest in some held out sample. So effectively we're going to fall back to a supervised learning [00:56:52] problem where within domain there's going to be what's effectively a verifiable common task in the words of David Donoho where we're going to ask are we producing hypotheses that explain variation in the outcome we're interested in studying. [00:57:05] So let me first very briefly talk about the scoring step. So the scoring step is going to try to construct a score that I'm going to denote by R of X which is going to associate each piece of unstructured data with a number or vector. [00:57:18] And what that's going to do is the score is going to turn unstructured data into something that we can now systematically compare. We can try to look for what are differences in the unstructured data that lead to changes in the score. [00:57:29] Uh there are many common choices of this so let me just highlight three of them that you see in the literature so far. [00:57:34] The first would be to define the score as the outcome itself. This is the simplest choice. If we're just trying to study a binary outcome, whether a defendant is detained or released, a case is opened or not, we could group unstructured data by whether it is a case note associated with an open case [00:57:48] or not. The second might be to define the score to be a learned predictor. Let's predict the outcome of interest using the unstructured data. So, for example, let's score each mugshot by how we uh uh it's predicted detention risk. Um and [00:58:02] the value of this is going to is going to be uh uh ripple through downstream in some choices that we can eventually make. Uh and the last choice you see would be to associate each piece of unstructured data, uh reduce it to an embedding. In other words, let's try to [00:58:15] place this unstructured object into a numeric vector space, where in that space, nearby points are going to, roughly speaking, express similar things. Uh and that's going to be, you know, has some familiar ancestors to ideas in in topic models. [00:58:30] So, how do you think about this choice? [00:58:33] Well, let's first think about the choice between the outcome and predictions of the outcome itself. I think in general, there's probably no reason to use the outcome directly. The smooth score is usually preferred. Uh we usually think of outcomes in these sort of social science settings as noisy, and trying to predict the outcome using the [00:58:47] unstructured data attempts to preserve only the variation in Y that is explainable by the unstructured data we observe. Uh and furthermore, this is going to open up avenues downstream when we do hypothesis generation, because having a predictor will allow us to score units that never appeared in the [00:59:02] data, which we can then exploit later on. [00:59:05] Uh an embedding sort of sits in a different category, because each unit is now associated with a point in a multi-dimensional space. So, what would be an example? Uh in Baron et al., we could try to take case notes, embed it in a space, uh an embedding space, where [00:59:19] we would hope that case notes that describe domestic violence are going to cluster together, far away from case notes that describe housing conditions that the child is experiencing. [00:59:28] Now, when you do this, each direction in this embedding space is now a potential score. We could try to ask how how much does a particular unit X align with some dimension in this latent space? So, that still sort of punts on the question because now you have to pick a direction [00:59:43] in this complicated space and then decide how to interpret it. [00:59:47] Uh one idea that you'll see in this broader space on machine learning and AI is that embeddings often behave more like an index. If you think about directions in an embedding space, they often bundle several concepts together [01:00:00] at the same time. People sometimes refer to this around the jargon of superposition and polysemanticity. [01:00:07] But as a result, when I talk about embeddings, I'm often going to refer to a class of models that people refer to as sparse autoencoders, which attempt to take some pretrained embedding and unbundle that index, re-express it in a sparse embedding space, and then [01:00:21] interpret that directly. But that's a detail we'll come back to later. [01:00:26] So, once you've picked a score, the next step in hypothesis generation procedure is what I'm going to call contrasting, which attempts to use the score to construct a set of examples E, units that differ along the score R of X. [01:00:40] The working presumption in this literature is that hypotheses can be labeled by inspecting differences. If we have examples that differ along the score X, we could pass that along to some human or an AI agent, and that would be easy for them to try to label [01:00:53] what are the differences between the objects in that set. Again, this is Polyani's paradox appearing. We can recognize hypotheses when we see them, but we can't ex ante generate them all. [01:01:04] Uh at this point, there are really three common strategies you see in the literature for producing these example sets, and each of them is going to return a different one. The first is what I'm going to refer to as morphs. [01:01:14] Let's try to generate a counterfactual pair unstructured data X and X prime by moving along the gradient of the score inside a generative model of the data. [01:01:23] Going to make that concrete in a little bit. The second is what I'm going to refer to as tails. Let's pick two real units in our sample from two tails of the score, high scoring units and low scoring units. That's going to give us two sets that we can kind of contrast. [01:01:37] And the last one is I'm going to refer to activations. Let's pick real units that are selected based on a sparse auto encoders embedding space. You know, units that are going to score high along one dimension in that embedding space. [01:01:50] As you can kind of see, each of these different contrasting strategy leans on a different choice of what's the scoring function. Morphs are going to require us to be able to score units that don't exist in the data. So, you need a predictor of the outcome given the unstructured data. [01:02:04] Tails are going to rank observed units and activations are also going to rank observed units. So, this is all very abstract. So, let me try to make this a bit more concrete by illustrating some of these choices in the examples that I mentioned. [01:02:17] So, first let's think about Ludwig and Melinathen. So, for them in their hypothesis generation procedure, the score is a convolutional neural network that they trained to predict detention decisions from the mugshot alone. [01:02:28] They're going to try to construct a contrasting set of examples through morphing. Let's construct pairs of mugshots that are close to one another but differ on the score. And again, the idea would be that if you just randomly sampled faces, they're going to differ in many different ways and it might be [01:02:42] hard to really inspect what are the differences. [01:02:46] So, how are they going to try to operationalize this morphing idea? Well, let's start with the simplest idea. [01:02:51] Let's take take a mugshot and just follow the scores gradient in pixel space. The scores gradient in pixel space is the direction that most increases predicted attention. So, that seems like a very reasonable place to start. [01:03:04] If you do that and implement it literally, you produce something that is decidedly not a face. So, we started with some synthetically generated mugshots, updated each pixel in this image along the predictions of our trained [01:03:17] convolutional network, clearly you end up with something that's a little more haunting than a photo of a face. [01:03:24] So, what went wrong? What went wrong is that when we think about unstructured data, this is in a really rich space. In the space of all possible images, you know, think of you know, 500 by some by 500 some matrices or tensors, in the [01:03:37] space of all possible images, real faces occupy an extraordinarily low dimensional manifold. And so, if you take any direction in that space, you're going to be stepping off this face manifold. [01:03:47] How could we fix that? Let's fix that by building a generative model of the mugshot distribution. In this paper, a generative adversarial network. [01:03:55] There's backup slides, there's a lot of resources on the internet to understand what those ideas are, but let's build some generative model of the mugshot distribution, and rather than morphing pixels, let's try to morph in the latent space of that generative model. [01:04:08] And so, the idea would be let's now think about the gradient of our score through the latent space of this generative adversarial network, update in that, and then use the generative model to decode. [01:04:19] And if you do that, you actually produce things that are sensible. So, as a sanity check, what Ludwig and Malina then report is they say, "Well, we want to interpret our detention predictor. [01:04:28] Let's start by predicting the age of the defendant, and then applying our morphing procedure to our age predictor." And if you do that, you actually produce morph pairs that allow subjects to pick out aging. So, they started with the same original mugshots, now applied this [01:04:42] morphing procedure to follow a convolutional network trained to predict age from the unstructured data, and produced a morph page morph face that looks like an aged version of the one we started with. [01:04:53] So, now they're going to take this idea and then apply it to their detention predictor, and that's going to produce many pairs of an initial mugshot and a morphed mugshot that differ based on the predictions of this convolutional network for who is more or less likely [01:05:07] to be detained. So, it's important to note that this morphing idea, nothing here is specific to faces in detention. We can apply this idea in principle wherever we have a good predictor and a good generative model, we can morph. In a wonderful [01:05:20] recent paper by Ziad Obermeyer and some co-authors, they apply the same technique to think about how could we extract features that are predictive of sudden cardiac death. And so to do so, they build a generative model of ECG waveforms and start morphing in the [01:05:34] latent space of that generative model using a sequence predictor to predict sudden cardiac death from ECGs. [01:05:41] So, images and waveforms have a good generative models, we can build those quite easily and those generative models have latent spaces that we can manipulate and sample from. So, you've heard about large language models which are generative models of text. How could we do this with text as well? [01:05:55] So, let's go to the next example by Baron and our team of co-authors. What we did in this paper is we scored each case note again by predicting case opening decisions using off-the-shelf pre-trained embeddings of the case notes. So, what if we try to reuse the [01:06:09] morphing recipe? What we would use a language model as the the generative model of text. [01:06:14] Um, there is a whole separate conversation we could have about why that is actually a very non-trivial problem in the mechanistic or interpretability literature. [01:06:23] Um, and I will just bracket that off, but will be happy to chat with folks about why it's actually very non-trivial to steer in the latent space of a language model and then sample from it and produce things that are are sensible. [01:06:35] So, instead, what did we do in this paper on case notes? Instead, what we did is we picked the second contrasting strategy, tails. We binned cases by predictions of case opening decisions using only the structured characteristics X. And then within each [01:06:50] bin, so case notes that have the same predicted probability of case opening decisions, if you were just using structured information, we asked, "What are case notes that are predicted to be very likely to have a case opening using just the text versus very unlikely to [01:07:03] have a case opening based on the text." And the idea would be this is a contrast set now where differences in the score are only reflective of differences in the text, not the structured characteristics that we already know and how did we understand how to interpret. [01:07:17] Um as our last example, let's come back to this reanalysis of inflation narratives. So, as I mentioned, this team of co-authors collected open-ended explanations of how households think about inflation. [01:07:27] So, as a starting point, you know, I did this in one day using Codex. This is very simple. You can do it, too. Uh converted each response from this paper into an off-the-shelf embedding. Just use OpenAI's off-the-shelf embedding models. And now once I have this [01:07:41] embedding space, if I ask a question, you know, what's a direction that we should choose? So, I'm going to pull from some recent work in computer science, which has a beautiful name, HypothaSAEs, um where I'm going to take the embeddings from my OpenAI model and produce a sparse autoencoder, which is [01:07:56] just going to re-express this embedding space to have many sparse features. [01:08:00] And now that gives me a candidate score, one per each dimension in this SAE's embedding space. [01:08:07] Um to then select which direction you're going to focus on, I'm going to split the sample into a train set, a validation set, and a test set. On the validation set, I'm going to examine which SAE directions are most correlated with inflation expectations, and those [01:08:20] are the ones I'm going to pick. And so then now my contrasting set are going to be the open-ended responses that score high versus low on the top 20 directions in the sparse autoencoder's embedding space. [01:08:34] So, once we've constructed these contrasting sets, we've now arrived at the point where we could try to start generating hypotheses. So, this is the third stage where I'm going to call interpretation. So, as I mentioned, contrast hands us concrete examples of unstructured data that differ along the [01:08:47] score. How do we go from these examples to hypotheses? So, how do people generate hypotheses? You know, perhaps you do it in the shower, but I think the way we all do it is we do it by searching for differences in the world. [01:08:59] We think about some outcome we want to explain. Let's think about observations that have high outcomes, observations that have low outcomes, and try to label the differences between them. And hypothesis generation procedures are just going to do the same thing using the contrasting sets that we just produced. [01:09:14] So, I'm going to call an interpreter iota, which is going to take the examples, and then return hypotheses. [01:09:19] And this is the stage that has been, you know, by far more than anything else changed by the, you know, rapid progress in AI and large language models over the last year and a half or so. We have now a really meaningful choice over who the interpreter is. [01:09:32] And so, existing work really uses two stages for this interpretation. One are humans. Perhaps we would recruit subjects online to expect the examples and name the differences they see. [01:09:41] That's what Ludwig and Melinathen did. In Obermeyer's analysis of ECGs, you can't just recruit people on Prolific to inspect differences in ECGs. You actually need to get doctors to look at them and name what are the more morphological differences. So, that's what they do. But in many other [01:09:56] examples, perhaps an off-the-shelf large language model could do this quite well. [01:10:00] And that's what we did in this work with Baron et al. and that's what we also are going to do in this reanalysis of inflation expectations. [01:10:08] This is I think is a fascinating question, how to think about the choice of different interpreters here. [01:10:14] Different interpreters, when given the same examples, need not return the same hypotheses. And if you go to, you know, all sorts of work in AI, we have ample reason to think that humans and large language models carry very different inductive biases in what they notice [01:10:27] from these sorts of examples. There's ample evidence that human and artificial intelligence really differ. [01:10:33] And here ideas and hypothesis generation connects to stuff that you may have heard about about language models having jagged frontiers of intelligence, humans being very poor at trying to predict what a language models will succeed or fail, or language models having their [01:10:47] own sort of systematic reasoning errors that look quite different than the sorts of errors that people make. [01:10:54] So, once we've gone through all this process, we have scored our unstructured data, we've constructed our contrasts, we provided them to interpreters, we now have a handful of hypotheses. What do we do? We come to the last step of evaluation. [01:11:08] So, everything so far returns a collection of hypotheses, which are features that are expressed in the unstructured data. And evaluation asks a very simple question, how predictive of the outcome are these named features [01:11:21] from unstructured data? And as a result, evaluation looks a whole lot like a very standard supervised learning evaluation problem. Let's have a holdout test set and just evaluate how well we can predict the outcome on the test set [01:11:34] using these generated hypotheses. And you know, really the core idea here is that if our hypothesis generation procedure never touches our holdout sample, it really does not matter to us how we produce hypotheses. You know, if things are nice and IID, if we randomly [01:11:48] hold out a set of hold hold out a set of examples, we can evaluate the performance of these hypotheses on this. [01:11:56] And as a result, you know, I really think about so far evaluation has been centered on this common task that's really powered a lot of breakthroughs in AI and ML over the last 15 years or so. [01:12:06] So, what do people report when they go on the test sets? You know, first they might report something that I'm going to call validity of each hypothesis, not to overload the word of validity in this this methods lecture. But at a high level, this tries to answer the question, how correlated is this feature [01:12:20] of the unstructured data with the outcome Y? [01:12:23] And so, if the outcome is binary, sorry, if the feature is binary, we might just try to ask, what is the average difference among units where the feature is equal to one relative to units where the feature is zero. [01:12:33] Whatever you end up choosing, this becomes a very standard sort of inference problem on the test set. [01:12:39] As Melissa Flag, you know, many of these procedures can produce a large number of candidates. And as a result, corrections for multiple hypothesis testing are essential. [01:12:48] So, in one paper in computer science that tries to use large language models for hypothesis generation, they have a pipeline that produces over 3,000 generated candidate hypotheses, and then when they go to the test set, they have to apply pretty significant multiple [01:13:01] hypothesis testing corrections, and only about 13 of them still remain significant. [01:13:06] And then there's been some more recent work in the econometric side thinking about this multiple testing problem. [01:13:12] Furthermore, the test set is not going to be the end point for us. We would like to select even the best performers on the test sets for further downstream analysis, and this is where problems in hypothesis generation, you know, perhaps connect back to ideas you've seen in econometrics on doing inference on [01:13:27] winners. We would like to select the most interesting, most promising hypothesis for downstream analysis, so standard ideas from the selective inference literature can be applied on the testing set as well. [01:13:38] A last benchmark people often report in this hypothesis generation space is what I'm going to call the completeness of the generated hypotheses. So, our procedures take unstructured data, try to distill them down to lower-dimensional structured features [01:13:53] that we could try to measure. So, it's valuable to ask how much of the signal in the original unstructured data have we preserved or do our hypotheses still capture? And the idea is actually to borrow an insight from work at the [01:14:06] intersection of economic theory and machine learning algorithm machine learning algorithms and calculate something that people refer to as completeness. [01:14:15] So, what's the idea here? Is let's build a benchmark predictor of the outcome from the unstructured data itself. So, in Ludwig and Mullainathan, that might be their original convolutional network to predict detention decisions from mugshots. [01:14:28] And then let's also build some naive predictor of the outcome, and let's ask how does the predictive performance of predicting the outcome using our structured hypotheses compare to the predictions the predictive performance of a model that just tries to flexibly [01:14:43] predict the outcome from the unstructured data itself. And ideally, our hypotheses would preserve most of the predictive signal from the unstructured data that we started with. [01:14:54] So, I really am focusing so far, largely because that's what the work that's been done so far is focused on on two sorts of evaluations. We could report on a held out test set validity and completeness. And I've really sort of taken the perspective that the [01:15:08] hypothesis generation procedure was applied on some separate training sample, and we're then just treating the hypotheses as some fixed object. But in principle, you know, I could have split the data in many ways, and so we could imagine resampling the data, doing [01:15:22] different splits, and now the validity and the completeness of the the hypotheses themselves are random, and their validity and completeness are going to be random as well. Um and so you could start imagine trying to build up ideas like an analog of size when we [01:15:35] think about hypothesis generation, or thinking of completeness not as a property of hypotheses, but as a property of hypothesis generation procedure. [01:15:43] Um and I think there is really interesting methodological questions on how you might think of whether there's a trade-off between size and completeness when we're building hypothesis generation procedures. Um but I will leave that for for the folks interested in these sorts of questions for later on. [01:15:58] And the last thing I'm just want to raise is I've also really only focused on an evaluation of how well do the hypotheses predict the outcome of interest. That's both validity and completeness. But those are not necessarily the only desiderata we would [01:16:12] want for the hypotheses we generate. For example, we might have a notion of wanting to generate novel hypotheses. [01:16:19] Then comes the question of novel relative to what? Um in some work that I've done with Sendhil Mullainathan, we thought about this problem in the context of how would we generate hypotheses that are novel relative to some structured parametric theory. [01:16:30] Um and then sort of for methodological folks, you might want to then ask, are there trade-offs across these desiderata? How do we think about what it means to build good hypothesis generation procedures? [01:16:41] So, in the last 5 minutes, I to now wrap this up by going back to the three examples and trying to show you what exactly do these procedures produce in these empirical settings. [01:16:51] So, let's go back to Ludwig and Mullainathan's analysis of pre-trial release decisions. So, at a high level, how are they going to partition hypothesis generation into these four steps? [01:17:01] They're going to have a predictor that predicts detention decisions using mugshots. That's going to be their score. [01:17:07] The way they're going to contrast, as I mentioned earlier, is they're going to build a generative model of mugshots and then start morphing that generative model along gradients in this detention prediction detention predictor. And that's going to [01:17:20] produce synthetic mugshot pairs. Then they're going to pass those synthetic mugshot pairs, you know, they probably did this in 2022 before language models were a big thing to some folks on Prolific to have them inspect the pair of mugshots and label what are the differences between them. And that's [01:17:34] going to be their interpreter. It's going to produce some hypotheses. [01:17:37] So, here is one such example of a morphed pair that the authors produced. [01:17:41] This is the same synthetic face moved from low to high predicted detention risk. So, what changed? When they presented this to the interpreters in their study, the interpreters overwhelmingly produced a hypothesis of being more well-groomed. [01:17:56] And it turns out that being more well-groomed is less likely to be detained, perhaps not the most surprising finding. But they can then play this procedure again, where now that they have generated the well-groomed hypothesis, let's now morph in a direction that's orthogonal to [01:18:10] well-groomedness and produce new morphed pairs of mugshots. So, this is again a synthetically generated mugshot and one that's now been morphed in the direction to have a higher predicted probability of being released sorry, a lower probability of being detained in a direction that's [01:18:24] orthogonal to well-groomedness in their get generative their GANs latent space. [01:18:29] This produces a second named hypothesis of defendants being more or less heavy-faced. And then when they go out of sample, it turns out that that is also enormously predictive of whether a defendant gets released or detained. So, now they have two hypotheses of facial [01:18:42] characteristics, well-groomingness and being or heavy-faced or not. And so, then the authors end by computing a completeness analog by pointing out that well-groomed, heavy-faced, and then all the other known facial features in the [01:18:55] psychological psychology literature together explain about 27% of the variation in their convolutional network. So, in principle, there may be more structure here that the authors are not exploring in this paper. [01:19:07] So, what do the authors then do once they have the hypotheses? They spend, you know, probably many, many pages in section six in the appendix to try to probe the novelty. They prevent a present a bunch of sort of more standard empirical evidence that these hypotheses are new to the psychology literature, [01:19:21] they're new to the crime literature, and they're new to practitioners, prosecutors, and judges. So, what's interesting here is this is a setting where we used unstructured data and hypothesis generation procedures to discover a potential mechanism for judicial error in pretrial release. And [01:19:35] as I mentioned, this method is general. Effectively, the same pipeline was applied to a ECG predictor of sudden cardiac death and actually produced a novel morpholo- morphological feature of ECGs that's correlated with sudden cardiac death, which is is [01:19:49] quite frankly an amazing finding. Um let's go to Baron et al. thinking about child protective services. Um there, the case opening predictor was our score. We constructed contrasts by binning case notes of those who are similar based on predicted case opening [01:20:03] using structured characteristics, but quite different based on embeddings of the case notes, and then we pass them to a language model to read and name them. [01:20:11] Um the LLM interpreters in this context would return five evidence themes that appear to consistently arise across case notes in a manner that's correlated with case opening decisions. Uh the most interesting of them are explicit [01:20:24] noticing of absence or ambiguity or contradiction of abuse evidence or noticing evidence that is exculpatory of exculpatory. I never say that quite right, exculpatory um for the family. [01:20:36] On a held-out sample, we can also report uh completeness measure. These five themes actually turn out to explain about 84% of the predictive signal for case notes uh for case opening decisions from the case notes. [01:20:48] And what's even more interesting is that there are systematic differences in what higher-performing investigators report relative to lower-performing investigators. Um and in particular, higher-performing investigators are much more likely to record in their case notes explicit reference of [01:21:01] corroborating that there's adequate care for the family or that there is ambiguous or contradictory evidence relative to the initial referral. So, it appears that the best investigators are reporting disconfirming evidence more often, which if you go to psychology is a classic corrective for confirmation [01:21:15] bias and has been really extensively used in some recent interventions to promote system two thinking in consequential settings. [01:21:23] Um very briefly, let me just show you the results from reanalyzing inflation expectations. Here, I'm just going to pull something off the shelf from the CS literature, that HypothesAE paper, and apply it to their data. [01:21:34] Uh what we produce are about 20 SAE features that add substantial signal over structured covariates for predicting inflation expectations. Um and then from those SAE features, we can pass them to a language model to produce 20 named concepts that are related to [01:21:49] inflation expectations, and those two add substantial signal for predicting um You can then try to correlate what are the AI-generated concepts relative to the authors' hand-coded concepts. Um and what you find is there is a large [01:22:02] fraction of the AI-generated concepts that really line up with what the authors hand-coded from their analysis of the open-ended responses. But there's a large subset of them that are not quite correlated. So, perhaps there is signal here um that is still interesting to explore. So, for example, one of the [01:22:17] themes that the language model generates based on the open-ended responses is whether a respondent explicitly refers to Joe Biden by nickname. That's very predictive of household uh inflation expectations, which you could think of as hyper partisanship. [01:22:31] So, to wrap up, you know, what I hope I convinced you just at a very high, very quick level is that we now have procedures that allow us to turn unstructured data into hypotheses. [01:22:41] What's going to happen next? These generated hypotheses are now the fodder for downstream empirical work. They may give us ideas about what are interesting mechanisms to test, what interventions we should trial, or what are causal effects that we may have overlooked. So, for example, in child protective [01:22:54] services, would training lower performing investigators to search for disconfirming evidence improve decisions, which is an idea that's been explored in recent work on police misconduct. [01:23:04] What I want to emphasize and leave you with is that the barrier to entry here is really low. I threw a lot at you in 40 minutes. The simplest version of this simply hands an LLM contrasting documents, asks it to name what differs, and then you the researcher inspect [01:23:18] whether those are interesting and try to validate them out of sample. More elaborate versions are actually off-the-shelf. The SAE framework I told you has really nice, good packages. [01:23:28] Yourself, Codex, or a very smart RA could do this themselves, and all of the embeddings I'm using are pre-trained off-the-shelf using API calls. What's really scarce here is actually your comparative advantage, which is finding economic settings where there is interesting unstructured data tied to an [01:23:43] outcome we care about. And if you have that, you're off to the races to engage in this activity. There are many methodological problems here for folks interested. I'd be really happy to chat with you about them. I think more econometrics in this space is really valuable. [01:23:58] And I want to flag that the frontier is moving rapidly. There is huge returns to deep interdisciplinary engagement with our colleagues in computer science and AI working on these problems. They have tools, they have ideas, they have algorithms, we have problems, and those two actually don't interface in a way [01:24:13] you might expect. So, to wrap up, once we have generated a hypothesis, we have to measure it. We have to turn this named concept into a variable that we can use across a full data stream and then plug those measurements for downstream estimates. [01:24:27] So, whether the hypothesis comes from a human or a machine, the next thing we're going to try to take up after the break is can we use AI to measure it at scale? [01:24:34] And that's where we're going to spend, you know, the last half of this methods lecture talking about. Um so, that's all I have for you. I think we have a break now until about 4:40, so 10 minutes or so and then someone will ding a bell. Um but [01:24:47] that's it. Enjoy the break. >> Mhm. [01:32:04] >> Mhm. >> Mhm.