Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # GPT As A Measurement Tool Authors: Hemanth Asirvatham, Elliott P. Mokski, Andrei Shleifer Discussant: Matthew O. Jackson Video: https://www.youtube.com/watch?v=VdvT0JzwMHU&t=10332s ## Talk (02:52:12 – 03:13:13) [02:52:28] Hello. Um, I'm Hmon. Uh, this is Elliot. Uh, that's Andre over there, but I'm pretty sure you could have figured those out. Um, we're going to shift a bit from uh talking about LLMs and AI modeling [02:52:42] human behavior to AI manifesting a specific human skill, specifically in the idea of comprehension. and what extraordinary things we can do in social science with comprehension at scale. So just a the [02:52:56] basic idea of what we want to get into first. What do we mean by this idea of GPT as a comprehension machine or as a measurement tool? Um sorry I'll keep saying GPT because I work at OpenAI but it applies to other things too. Um some examples of this in operation what how [02:53:10] this can actually show up in real research. um validation of how this can actually be um confirmed and that it actually works as we intend and then a new application to the history of technology adoption and learning some [02:53:22] lessons from that. The big idea we want to start with is that in the social science, in economics, we really care about, you know, what people do, how people behave. But almost all that, so much of that data is encoded [02:53:35] qualitatively. It's in text, it's in blog posts, and diaries, in websites and curricula. It's in the entire history of the NBR slide decks and transcripts and all of that sort of stuff. And so there's all of that data, and we kind of [02:53:50] have two choices at social sciences. There's like the sociologist route which is like I'm going to really interview a few people interview like look at some data really really closely and understand the richness of it at that qualitative scale with human intelligence and comprehension and I'm [02:54:04] going to do it at a low end so you can't have that statistical rigor or I'm going to go down the more econ route of I'm going to use the sliver of data that is quantitative that is abstracted unfortunately but allows us to do analysis at bigger scales and get [02:54:19] statistical rigor and it's really been choice of those two routes. And our proposition here is that with GPT, there doesn't really need to be that choice anymore because you can have something that can make comprehension do that sort of um task of reading the richness of [02:54:32] data, but do it at extraordinary scales and at very low costs. [02:54:37] So what we bring to the table with this paper is we I mean we have a specific package we built like a a open source Python package um called Gabriel. Um, but just to be clear, everything we're talking about here just talk is about generally about LLMs, generally about [02:54:51] AI. This is just a tool to help anybody use it. What Gabriel is doing, let's I'll get it has many functions. I'll just give you one idea. Like say we have some qualitative data and we want to ask some questions about it. We want to measure something about it. Um, say for [02:55:06] example, it's the state of the unions. We could take a state of the union. Um, we can ask like okay on a zero to 100 scale, you could ask this to catch right now. How populist is this? How individual is this? How like any sort of complex human concept? Just ask that [02:55:20] question as you might a human labeler and then see what GPT says. Now this is this is this idea entire idea. It's not um you know tuning a model. It's not technically complicated. It's just [02:55:33] asking that sort of comprehension question to the model and seeing um what it says. So you know the speech about the space program will be very much about tech optimism and get a high rating for that. Um the speech about [02:55:47] indiv like about um American rugged individualism will get a rating on something else. [02:55:52] So the idea at play is that that very simple thing you can do with chat GBT. [02:55:57] The really interesting thing is what you can do at scale. And the key shift here from human labeling because we could do all this stuff before is just how different the price and the um experience of and scalability of this is like you know with a h 100,000 full text [02:56:12] church sermons that's on the order of a million dollars to rate 10 attributes 10 you know human conceptual attributes with human labelers it's 50 bucks with GPT. Um, so the scale is really important here, not just because it's [02:56:26] way cheaper and more accessible, but because there's a whole host of new research questions that you can answer now or you can try to ask because it's so much cheaper, because it's so much more feasible. [02:56:37] So just to be clear again, what we're going for here is something that is the lessons we're talking about are broadly applicable to GPT. The stuff we put forward in our paper with Gabriel with our package, I I think the best analogy is like STA. Anybody can code a linear [02:56:51] regression model. In fact, with like modern coding tools like Codex or cloud code or whatever, it's pretty easy to do it. But people still migrate to stuff like STA because it's good to have a standardized source of measurement. It's good to have some sort of reference validity and something that's [02:57:05] consistently used. And there's more complicated methods than the most basic ones like the Gabriel package has and we're happy to talk about. Um, but we're going to start with the really simple thing which is just, hey GPT, how populous is this speech 0 to 100? [02:57:20] So let's go through some examples of applying this in practice. First, basically the entire congressional record. So let's just take from 1880 to the near the present. We're gonna take all the speeches made on the Congress [02:57:33] floor um by one representative in one year, smoosh them together, and just for one of those representative year transcripts, we're going to ask GPT um a bunch of questions about how pro- foreign how pro- foreign intervention is [02:57:47] this congress person in this year? How um tech optimistic is this person? How much anger towards the other party does this person do? And because it's so cheap, we can do this across an incredible amount of attributes really quite easily. here. This one shows, I think, a pretty interesting trend, like [02:58:02] following some things we'd expect, like Congress becomes, you know, much more pro- foreign interventionism, you know, following World War II. We see other patterns in that if you look closer. But I think the more striking thing is just how completely bipartisan it is. To be [02:58:15] clear, again, it's GPT is only looking at one of these representative speeches sets at a time. It's not seeing this overall trend. It's not trying to do some big guessing game. It's looking at only one of these confined things. And yet it still seems to not be able to [02:58:28] distinguish Democrats from Republicans the vast majority of the time on average. We actually see a somewhat similar story on moral universalism and a whole set of progressivecoded attributes um where um you see differences between Democrats and [02:58:42] Republicans. They varying over time. In fact, in the '9s and the early 2000s, you actually see them basically being identical, like they are indistinguishable on moral universalism and a bunch of other um progressive attributes. And then suddenly there's a [02:58:56] divergence in 2008. Um so it it we were able to see this with granularity. Um a different story of partisanship at least in Congress where you see not a gradual shift but a pretty sudden shift in 2008 from an era of at least in this data the [02:59:11] most bipartisan era of uh Democrats and Republicans on these progressive attributes. [02:59:17] And another one we can see optimism about technology increasing over time. [02:59:21] you know, contrary to some narratives about, you know, a little bit of a doldrum we're in right now in how people think of technology over history, that narrative might not be as simple as as we think with some of those current explanations. Switching gears to another [02:59:35] um category of analysis. This is an analyzing the question of how toxic is social media, how toxic is the internet. [02:59:41] Again, with the granularity and the scale we can do these analysis, we basically take a bunch of subreddits. I don't know how many of you have read it, but basically there is these internet communities like you know there's the little Pokemon forum where all the people talk about Pokemon and the one [02:59:54] where they talk about politics and all sorts of other things. And we basically for each community we take a bunch of the threads, a bunch of those conversations, measure toxicity, measure admitting you're wrong, measuring a bunch of these complex human concepts [03:00:08] and we analyze those things. For toxicity there is a decent amount of toxicity on average but it's extremely heterogeneous. There's a few communities where you see a ton of it and then a bunch of communities where you see none of it. In fact, we show that basically the communities you'd expect kids to be [03:00:23] in have way less toxicity. So the narratives around what kids are exposed to and don't just match the big picture narratives we could show before this sort of um granular measurement. [03:00:34] Um our third example we'll talk about here is on school curricula. So we've these examples we just talked about are about GPT analyzing data sets data text data. we have this is about GPT also collecting the data using its um prowess [03:00:48] as a web searcher in the last year or so um that it's gained. So basically the steps are fairly simple. You start GPT like you would do with chat GPT one instance of GPT one county in the US. [03:00:59] Hey GPT, go research like an RA everything you can find about this county um about how US history is taught in schools in this county and um rate um and then write these long reports about [03:01:13] what is taught in those things. Find syllabi, find homework assignments, find everything from school districts in that county and then write these reports and then do these pair-wise comparisons is another method in Gabriel of between county A's report, county B's report [03:01:28] across the data set. If you play chess, you can create these ELO like rankings around out of that and create these ordered rankings for how much each county focuses on different things in their US history curricula. And so, for example, you can see that focus on [03:01:41] historical racism is, as you might expect, much higher in um more liberal areas, but it's if you control for liberalness, you actually find that race doesn't really have much of a explanatory power there. mostly [03:01:54] explained by political lean and some things related to um um like um education and stuff like that. Um a different completely different pattern if you look at things like focus on rugged individualism. You see that as you might expect in the rural west in [03:02:07] Texas in places like that. some more surprising patterns being pro technology, the f focus on technology and historical innovation, highest in the south of the US in these high school US history curricula and um even more [03:02:21] surprisingly to me at least the positive portrayal of the New Deal and you actually kind of see Tennessee light up on there um kind of like um the history might have indicated um in retrospect. [03:02:33] Um, so I'm going to talk a little bit now about validation, which I think is a really important component of this method. So, you know, the obvious question is we're presenting various things with LLM classifiers, LLM ratings, and we want to know that it's [03:02:46] working. And by working, we care about probably two things. The first, is it accurate? And by that I mean, does it match either human judgment um or some external ground truth? And it's not necessarily sensitive to specific prompts like what we built. And [03:03:01] secondly, is it direct? And by that I mean, are we measuring the variable that we say we're measuring or are we measuring some proxy? An example of that would be if I give a speech and I'm trying to measure how pro- environment [03:03:14] that speech is, am I really measuring whether it's pro- environment or is the model instead guessing and it's measuring whether it's liberal? That would be bad. Um, and so we've tried to be thoughtful about this validation and really understand the capabilities of [03:03:27] the models. And we have a series of tests um in the paper. The first starts by looking at a couple of well-known um data sets in a variety of sectors where there's human labels. For instance, the V democracy data set which looks at quality of institutions or Metacritic [03:03:42] movie reviews. And we find that you know the model matches human judgment extremely well. Um but one of the things that we try to do to get beyond just a few examples uh which might you know you might worry that these were selected or that performance is uneven across the [03:03:57] task. um we run a very exhaustive evaluation of performance across many many gold standard human text data sets uh which are taken off of hugging face. [03:04:06] And so we take every single data set on hugging face that has gold standard human labels which is more than 300 data sets in the end. Um and each of these have you know hundreds of observations or thousands of observations and we [03:04:19] replicate them using Gabriel and we look at performance across the board. Um and what we find is that you know the the performance is very high. Uh in general we find you know um this is true across [03:04:33] models not just for the most expensive models uh but also for for cheaper distilled models which we think is important for researchers meaning you know that this is truly a thing that is available at low cost. Um and this is true you know accuracy really holds [03:04:46] across the board. Um but it's not quite 100%. And so the question that we ask is is this a product of noise or is this a product of the models not being quite there? And the way that we get at this [03:04:58] is by comparing to human um inter you know like if multiple humans rated the same text uh the human variance how would the LLM compare to that and we look at data sets where there were individual human raiders of the same [03:05:12] observation and we compare with the models evaluation and we find that across the board the LLM is typically statistically indistinguishable from individual human raiders relative to the distribution and often better than the marginal human raider at matching the [03:05:26] human mean. um which gives some credence that you know this variance or the the noise that we're seeing is really just noise that this is a subjective task and it's hard to get uh perfect accuracy. We look at prompt wording and robustness and we find that for our purposes in [03:05:41] this specific use case of trying to understand uh unstructured data and qualitative data the prompts do not matter that much as in we we generate many many variants of the prompts that we use in the package and we find that in fact the accuracy holds quite well [03:05:55] across the board. Um, and this is to allay the fear, you know, that our prompts are some kind of magic. We don't think that that's true. Um, and we think that this is more a product of general comprehension. Um, we look at the possibility that performance on our [03:06:10] tests is a result of memorization as in that the models have read these labels and therefore there's high accuracy as a product of the models knowing what they're seeing. And so we in this hugging face um you know data set we [03:06:22] look at data sets published pre and post uh training cutoff and we find that there is in fact no difference in performance. Um and lastly we look a little bit at this idea that I referred to previously of can you measure whether the model is actually looking at the [03:06:37] variable that you want or at a proxy. And so on that example I cited of you know um envir pro- environment speeches um versus pro- liberal or like politically liberal speeches, we generate some speeches and then we look [03:06:51] at measuring both speeches with liberal content and liberal content removed. Um and we look at the pro-environmentalism of them and we find that there's in fact no difference which um gives us confidence that the model is measuring what we want. And we also do this in a [03:07:05] more general way by looking at uh regulation throughout the country using the web scraping approach and we find that you know when we remove irrelevant signals um the the scores are essentially the same. [03:07:17] >> Yeah. Um it was a real blow when I learned that uh our prompts were not magic but I think it's probably good for the research community. Um so now we're going to talk about um technology adoption and using GPT for a bigger um research demonstration using Gabriel. So [03:07:32] what we want to do the question we care about is understanding the history of technology adoption. How long does technologies take to adopt um how um other what how else can we characterize the history of technology? Now there are data sets that have been assembled [03:07:44] before on the order of maybe 100 200 maybe 300 technologies. But with GPT we can scale this to extraordinarily larger sizes by an order of magnitude. We just start with the entirety of Wikipedia 18 million titles you know including like [03:07:58] Bon Joy or whatever. Um, and we just use various methods in Gabriel to filter that down to historically significant technologies. So, you know, filtering, dduplicating, and a set of things like that. And we [03:08:11] get 37,000 historically significant technologies. And then we pull a bunch of attributes about them. When were they invented? When were they adopted by their target audience? Where were they invented? Who invented them? What institution? Um, what how large and [03:08:25] bulky were they were they? How much do they require highly specialized training to use? Um what category there were all these sorts of attributes. Again if GPT can comprehend and use its understanding can we leverage that to get a bunch of [03:08:38] useful data to do it in structured and statistically rigorous ways at scale. So big headline finding was that we can show that um the adoption lags of technology but we can show with granularity how this has been declining [03:08:52] secularly over the last um two centuries um from about 50 years in early industrialization from the invention to the wide adoption of technology by its target audience to closer to five years [03:09:04] today. Um and one variation in this is the category of technology. Now, our original hypothesis was something more softwarebased or more like AI or even um would be something you'd expect to be particularly fast because it's updatable [03:09:18] or easy to shift. But what we actually find is software is fast, but it's not fast because it's unique. It's because it's just modern that it's not relatively fast relative to its contemporaries. So, we do this by considering excess adoption lags. How much faster are you than your [03:09:32] contemporaries? We actually find military technology to really be quite fast there. Um and I think what we saw in the data was kind of technologies invented near the beginnings of wars needing to be adopted before the end of [03:09:42] the war to be useful. Um so um um so like people should be less critical of Boeing by that account. Um and then um we can also like just look at a bunch of trends over time of how technology has [03:09:55] evolved. So like because these are you know complicated attributes human um level attributes we can see you know that technology has gotten smaller um kind of from the 50s onwards. We can see that there was a steady increase over [03:10:08] time in the um in the requirement of specialized training to use the technologies that went up and up and up until the 50s and then with the computer age it's it's gone down um since then. [03:10:18] Um and we can see how many how these attributes explain shorter or longer lags in the technologies like needing network effects makes technologies adopted slower. Um you academic research sorry guys makes um technologies adopted [03:10:31] slower. Um but there's we we go through a bunch of attributes to explain variation in adoption lags. Um other things where was it invented? Um we can map the history of the United States rise as an invention powerhouse but also are able to capture I mean other [03:10:46] countries but able to capture how high that share is at least in our data set still of US invention dominance um even relative to its GDP share which is um quite a bit smaller. Um similarly we are able to do that at a subnational level [03:10:59] because we get the specific information with GPT doing all this research and using all of its knowledge. Um and we're able to show that like in US history three states um California, Massachusetts, New York account for more [03:11:12] in invention in US history than the other 47 put together. Um and the where these technologies come from. Um we see the rise of corporations um as being um [03:11:24] sources of technologies. um and uh and um the decrease of the classic early industrial independent inventor. Um so like you know Eli Whitney then um open AI hopefully today. Um so what does this [03:11:38] all say about AI? Um this was a part of our motivation for studying tech adoption. What it says about AI is actually the the things about the specific attributes of AI. It's kind of a mixed bag about whether it'll explain [03:11:52] whether AI gets adopted faster or slower. The first order factor is just that modern technologies are adopted much faster. So you know the old narratives about older generative purpose general purpose technologies don't necessarily apply to how we should think of tech adoption today because [03:12:06] things seem to be much faster overall like the core idea at play here is like yes chat GPT or codeex or whatever can be your RA. It can help you with coding. [03:12:16] It can help you with writing. It can help you a lot as a researcher. But what we want to profer is what can you do not with one GPT but with a million. What can you do when you have an extraordinary scale of cheap [03:12:30] intelligence at your at your behest? And so we think that just the the this whole realm of qualitative data, this whole realm of new information is opened up to us and it's all just all about us asking the models to do these sorts of tasks. [03:12:43] >> Thanks. >> Thank you. [applause] And our discussion, Matt Jackson. ## Discussant remarks (03:13:13 – 03:28:17) *Shared across the three papers in this session; Matthew O. Jackson discussed all three.* [03:13:13] Great. So thanks to Eric and Karen for organizing this and to assigning me to three very exciting papers to discuss. [03:13:22] So I want to start just with this picture and to I I don't know how many of you know historians but historians a lot of what they trained themselves on was collecting data going through [03:13:35] archives spending time making sure they could translate qualitative things or undigitized things into things that could be eventually used and curating it and making sure that they could somehow validate it and then going on to to [03:13:50] using it to to do some studies. And so one question is, you know, how long did it take to do this compared to how long did it take historians to to spend time collecting this kind of information? It couldn't be done at the scale that was [03:14:04] was done here in in a matter of probably uh days or weeks um to put this together. So it's very impressive and the scale is eye opening. So AI aided [03:14:17] research is going to be lead to an explosion in the quantity um and the potential for quality of research. And so um we need ways to harness it and evaluate it. And I think that that's you [03:14:30] know something that comes through in all of these is making sure that we're understanding what we're getting out of it. And so I want to talk um in a little bit of detail about this. And I'll start with just a picture uh from a paper on AI behavioral science that puts this in [03:14:45] context and a number of the authors on this are are in the room. Um so this is from a a workshop last year and it sort of breaks things down into three different categories. So we can think of assessing AI behavior. So AI is out [03:14:59] there. It's it's going to be being used in increasingly um complex ways. We want to understand how it's behaved. That's sort of the alignment problem. Um, we're going to be using AI for behavioral science. So, uh, using it to to actually [03:15:13] understand behavior and to to study humans. And then there's also understanding human AI interactions. And that was more what we saw in the first session today. And so this these three [03:15:27] papers fall mainly in this category of of using AI for behavioral sciences. [03:15:33] And you know when we think about why AI is useful, it's useful to sort of think about what it expands in terms of why is it better than humans at different kinds of tasks. And I think in in this in these papers, what we see is it can [03:15:47] digest a lot more data. It can make predictions. It's going to be useful in simulations and we can design and control it. So these are sort of different ways in which AI is enhancing human behavior and it's useful to keep those in mind when we're trying to [03:16:01] evaluate this. Okay. So so AI enhanced research um it can enhance the simulation techniques we have. It can help us develop new modeling tools. It's going to allow us [03:16:14] to do larger, more complex, better trained and calibrated simulations. It's going to be much cheaper than humans. Um it can run many more experiments than we can. It can infer things from prompts [03:16:26] and uh try and we we can use that to understand behavior and I'll talk more about that. And then it also has this possibility for new data analytics that we saw in the last paper where it can quantify qualitative data. It can [03:16:40] digitize things. We have ways of of doing visualizations and categorizations that we didn't have before. So there's a whole series of ways in which it's going to be enhancing. And that's not to mention the way that we use it every day, which is just, you know, doing if you're a theorist, you do proofs and [03:16:55] conjectures or coding. You can do lit reviews. It can do a whole series of things. So, it's it's going to be coming into our lives increasingly as scientists, researchers, people developing engineers, etc. Um, the [03:17:08] papers that we saw in this session, I think it's useful to sort of take the high level and think about where do they fit? They fit in these categories. We're going to be building larger, more complex, better trained simulators and [03:17:21] models of behavior ways in which we can um emulate humans or imitate humans, I should say. Um we can infer from the prompts, sorry, we can infer from the prompts and the settings about these [03:17:34] induced behaviors. So, uh I I'll talk about that especially with with respect to to Ben's paper. There'll be things we'll be able to infer about theories from this. And then we can also quantify qualitative data and use it in ways that [03:17:48] we were not able to before. And so what do each of these papers teach us? So with three papers, it's hard to sort of spend a lot of time on each paper. So I'm going to try to try and pull out big points from each of these papers and think about what are we learning more [03:18:02] generally. And one thing I want to emphasize is that, you know, LLMs are temporary. So we we're we're we're looking at this sort of particular object that's come along in the in the last couple of years. Five years from now, we'll probably be talking about a [03:18:16] different form of of AI or an enhanced version which is going to evolve in in ways that we can't anticipate at this point in time. So I want to keep this at a level of you know what do we learn in general rather than what do we learn specifically about these instantiations [03:18:31] of AI that are are you know quickly um becoming ubiquitous. So um Ben's paper made clear that we can begin to harness the steerability of LLMs and more [03:18:44] generally steerability of any kind of system and we can train it in some settings and then see how it performs in other settings and that's going to be a general you know it used to be agent-based modeling now it's going to [03:18:57] be um LLM based modeling or or other kinds of models. The advantage of that is that it it allows us to see exactly what we're putting into the models in ways that that we can build prompts that [03:19:11] generalize and try and understand what we're inferring from that. And I want to point one point out one thing that I think is important in terms of of the of of Ben and John's paper is that when you think about the the example that they [03:19:26] did with the 1120 game, they were using level K reasoning. So level K reasoning allows you to produce a human behavior map that's very nice. Now you could have used some other model instead of you [03:19:39] know that comes out. So suppose instead we want to test whether humans are risk averse or whether they're fair or altruistic. You could have instead tried to put that into that model and see whether it would have done it well in terms of this you know predicting the [03:19:53] 1120 game and it probably would have failed miserably. And so so that allows us to say actually there's something about the depth of strategic reasoning that is present in these kinds of settings. And so now this is a tool for [03:20:07] saying you know if you want to understand how people are behaving in certain kinds of strategic settings where they're interacting with other people and trying to anticipate them you need a model like this. So it allows us to to sort of you know go through and [03:20:21] decipher as I put in the bottom decipher human behaviors and as you change the game you change the models you're using we can learn a lot about that. So I think that it's can be a very powerful [03:20:32] tool for not only simulating but also you know going back and figuring out how do we understand human behavior by what we had to do to train the LLM in order to you know give a certain kind of [03:20:46] behavior out. So it can be a very useful tool in that manner. [03:20:51] Um, and I I think you know that gets back to the to the lab rats that that um Abishek mentioned and and effectively we'll be building different kinds of lab rats and we actually learn a lot about human disease by figuring out what kinds [03:21:06] of lab rats did we have to build in order to test a medicine and so so there's a a nice you know feedback there and and in terms of um Abishek's talk I think you know again this is teaching us [03:21:20] something about the harnessability and the steerability of LLMs. It goes a little more into the to the black box of LLMs and I want to come back to that point in a minute. But using these [03:21:32] sparse autoenccorders, probes prompts to steer in increasingly interpretable ways. And and that interpretation word was mentioned a bunch of times. And I I'll say a few things about interpretability. It's going to be [03:21:46] important in terms of us understanding how these things are working. And as AI becomes increasingly complex and opaque, it's more and more important for us to be able to understand why is it doing [03:22:01] what it's doing and how do we interpret it. And um it's useful to you know the the the idea of working on llama rather than other LLMs allows you to work with with the um sparse autoenccoders and the [03:22:14] probes in ways that you can't with other AI and and having access to the understanding of how these systems are working is going to become increasingly important. So, not just building better simulators, but understanding why [03:22:29] they're simulating in the ways that they are. And and that's going to be an issue. Um, in terms of Gabriel and the the the the paper um on on using this as a a systematic LLM tool to measure [03:22:44] qualitative data. Again, I think this points to both the enormous potential of us being able to replace a lot of wrote [03:22:56] um repeatable tasks of, you know, quantifying data, collecting it, analyzing it at at scale in ways that are going to be consistent. You know, if you employ four different RAS to go and do this and they each have a slightly [03:23:10] different encoding techniques and so forth, you might end up with data that's much worse than using a very simple um system that you can give explicit directions and you know it's following very specific rules in how it's coding [03:23:24] things. So it can do it with accuracy, it can do it directly as was mentioned and the cost replicability and so forth the scale is is unmatched. Uh I think one thing that that comes up in sort of [03:23:37] thinking about this is this approach could be used to start dealing with a problem that we've had for years which is understanding measurement error out there in data data sets that are existing and have been used quite a bit. [03:23:51] Um you know this this morning's talk about building everything on one pillar. [03:23:55] um understanding to what extent are we sort of using data which was collected by humans or constructed by humans in ways that might be somewhat idiosyncratic. So are there ways that [03:24:08] that we can use this kind of tool to better understand and collect parallel data sets um other kinds of data sets and begin to to understand measurement error better. So I think there's a lot of promise in this that goes beyond just [03:24:22] using it as a single tool for single studies but but using it in more general ways to understand measurement. [03:24:30] Okay, I want to say a few things that that I think are are sort of interesting big picture um issues that we face as AI enters into our world of of data and and so forth. And one is a lesson from a [03:24:44] data boom that came in the 1980s in finance and people you know there was suddenly minute-by-minut stock data that was available. People started analyzing it like crazy. We found the day of the week effects, the January effect, the [03:24:56] index effect, um large cap effect. There were all kinds of of particular anomalies that were coming up in sort of shiny new facts. And as we start deploying this um these tools, it's [03:25:09] possible that we can become uh you know just inundated with different interesting relationships that we couldn't have seen before. And I think it's that means that we have to start disciplining ourselves or we're just [03:25:22] going to be under a sea of of potential research questions. And so somehow managing to to understand what we should be asking, how we should we be asking it rather than just, you know, throwing it out at everything we can see is is going [03:25:36] to be hard and theory in the in in this age of AI can still be very useful in sort of guiding the deployment. What are the questions we should be asking? How do we assess um AI's use and and output? [03:25:49] Um and and how do we enrich the models in ways that are useful? [03:25:54] And I think there's um one kind of interesting paradox I've been thinking about quite a bit and and I think that this is going to become increasingly um important and and you might think of this as like the chess go paradox. So [03:26:08] it's been just about 30 years now since uh Deep Blue beat the a grandmaster. Um it's been about 10 years since uh a Go Alpha Go um beat one of the best Go players in the world. So, you know, [03:26:22] we're getting to a point where AI can do things in some ways better than humans, at least in very specific kinds of tasks. Then the question becomes at some point if it's doing things that we can't understand. So, it it generates kinds of [03:26:37] plays that look very peculiar and so forth. So, that teaches us things, but at the same time, it might be that we can, you know, it's it's doing things that are beyond our abilities to to [03:26:48] comprehend. Then how do we make use of that? Do we just trust it? So if this is, you know, if we get a simulator that does better at weather prediction, that's fine. Maybe we don't need to know exactly how why it's predicting exactly what time it's going to rain tomorrow [03:27:03] and how much it's going to rain in different parts of California. Um, that would be fine. Um, what if we start using it for predicting what's going to happen to three different potential rate decisions by the Fed and we say, "Okay, [03:27:17] here's it can give you really accurate predictions. We can't tell you why it's predicting that, but we can tell you it's doing that. So, these systems are going to be do able to do that. Um, do we want to understand why a certain rate would would lead to one thing in the [03:27:31] economy versus another? Um, and you know, eventually when trade policy, you know, understanding things like trade policy and what the implications of trade policy are, we can be using these tools to create new simulations, to [03:27:44] analyze larger and larger data sets, to build these systems. And these systems are going to become increasingly opaque. [03:27:50] And so the kinds of tools that we're talked about today to measure, understand, and again this this word mechanic or you know term mechanical interpretability, we need to be able to look under the hood eventually if we [03:28:04] want to understand the and explain why things are working. And so that I think this an interesting paradox that we're going to have that it's going to be doing some things we'll never be able to interpret and other things that we will and you know how do we handle that ## General Q&A (03:28:17 – 03:41:40) *Shared across the three papers in this session.* [03:28:17] trade-off. So thanks [applause] >> really interesting set of questions from from Matt and from the presenters. So why don't we line up over here for questions and uh while you're all doing [03:28:31] that I'll take the uh moderator's privilege to to ask the first first question but go ahead and line up there. [03:28:37] Um, so it's interesting all these papers thinking about where these things are seem incredibly powerful in many ways. [03:28:44] Where does the power come from? I was disappointed to learn that it's not magic after all, but it seems to come from the the knowledge that's in it. And I think there's some a bad aspect to that and a good aspect. One that I'd love to have people discuss a little bit more is the risk of contamination. I was [03:28:58] actually surprised at how bad the uh U u LLM did just out of the box on 1120 because the answers were right in the corpus and they could have just looked it up but but you know and sometimes they do that and you don't realize it. [03:29:11] Um, but also there's there's some potential for it to be really powerful. [03:29:15] And that also has an interesting thing is as as the technologies get more and more powerful, they're going to read all of your papers as well and learn the four steps to making a good uh simulator and know the relevant theories that it [03:29:29] should be drawing on, including some that we didn't think of in all likelihood. And then you sort of start wondering at at some point will the will the um out of the box version actually be so much better because it's it's drawing on everything and then more that that any of us can think about. So I'd [03:29:43] be curious how how you react to to those pros and cons. [03:29:48] And it's that's a question for any one of the the papers who presenters who like to Yeah. Go ahead. [03:29:56] >> Is this working? Um I think it's a great question. I'd first of all say like I'd love for my paper to be obsolete. I think that would be kind of cool. Uh I don't have a perfect answer, but I think there's a thing I've been thinking about like good [03:30:09] memorization and bad memorization for these foundation models for making predictions. And I think since it's an economist crowd, I think that good memor like bad memorization would be a student who is learning game theory for the [03:30:23] first time and memorizes the fact that defect effect is the pure strategy Nash that you should always play. like good memorization would be a student memorizing that a Nash equilibrium is when everyone is best responding to one [03:30:37] another. And I don't have a perfect answer of how I would differentiate between how to learn those two different types of memorization. I think they are in effect both memorization. Maybe like a cognitive scientist will argue with me that like one is a learning [03:30:51] memorization. I don't know how how to articulate those differences. But I think us being able to specify when models, us being able to understand those two different types of memorization better is going to be is going to be key. And I I've been thinking about it a lot. I don't have a [03:31:05] perfect answer, but I think that's the distinction. [snorts] I think that there is a difference between memorization and contamination [03:31:19] in the sense that just because a model knows something doesn't mean it actually is using it in making a decision or something like that. Um so like just because it knows of a game or knows a story um or even knows the outcome of something like we do some tests in our appendix of our paper on this like [03:31:33] doesn't mean it uses that when we're asking it do a measurement task. So I think you have to be like more careful if that like you know in a prediction task for example where that's um more likely to happen um on prior data. But um but I think that like part of what we [03:31:47] see in our paper is that like even though it could have known some of these labels or known the game or known stuff like that that um the models are you know kind of like we are capable of compartmentalizing to a certain extent. [03:31:57] We don't understand that full extent but but I think that's there. Um just to quickly add to that um in in some past work we've uh we've like actually looked for memorization when we know that the model was actually trained on the data and even then those rates of [03:32:12] memorization are maybe in the 10 to 20%. And part so I was I have I think one should be worried about memorization in a statistical way. It's kind of a case by case uh case by case question. But I think the the rubric that like just [03:32:27] because a model is trained on something is necessarily spitting things out verbatim from the text doesn't really hold partially because these models are trying to compress all the information that they're trained on. And so even if they're trained on a bunch of things, I [03:32:40] don't remember like chapter six of a book I read even though I read it in if I read that in third grade. So the models have kind of that same feature. [03:32:48] And so we have to start to get at a prompt by prompt or a case by case basis of what are they relying on when responding to a particular answer and that gets a bit bit more trickier. [03:32:58] >> Questions. >> Okay. So I have two questions for Benjamin. Um so first I wanted to know if you've tried any of the new models with your methodology like if you use GPT.5.4 four or 5.5 and if the baseline [03:33:12] improves when you use a better model and if the difference between your method and the baseline changes if the model improves and then the second question I was wondering on so you talked about the first moment on the distributions and that it it gets better I was wondering [03:33:26] on the second moment because we were talking before about kind of the the variance on the responses and how there's less variance for the interviews for example and I was wondering if the if the responses be if there's as much variance as uh there is for human [03:33:41] responses on on the AI responses or or does the variant shrink or uh if there's any comment on that. [03:33:49] >> So for the first question, I haven't tried more advanced models on our games specifically, but I have tried more advanced models doing the approach in other settings and I know other people have too. And I've I don't see a clear pattern that [03:34:03] the LLMs do a better job of predicting human responses unless human responses are like more rational. In that case, better models do tend to predict it. So, but in games like the 1120 game where it's hard to say what the most rational [03:34:18] thing is because the rational thing to do is to know how your opponent what your opponent is. Like I am rational to pick 18 if I know my opponent is a 19 picker. Um, so to that point, not not necessarily seeing it better across [03:34:32] settings and have tried the approach, although not for the specific games on other models. And then for the variance, uh, I haven't explicitly checked on the variance. I think you could kind of extend this to trying to match as many moments as possible as part of the procedure. I think that'd be like a cool [03:34:46] extension. And I think kind of the kind of what I would like to take what I would hope I would be conveying from the method is not that like one had to follow a particular optimization approach when trying to set up these uh simulations. I'm a little agnostic to [03:35:00] the optimization approach. I think there's lots of things you could try. [03:35:03] I'm guessing the more data you have, the more moments you could match probably means better, but it's an empirical question we'll have to explore. [03:35:11] >> Thanks. Okay, this is a question um about the application of a Gabriel. It's really interesting application we have here. My question is uh I want to take the author [03:35:23] stance on to what extent the post application human involvement are needed after you assemble the data. To take an example you know when we write papers there is a first stage um you know we circulate within the department in brown [03:35:38] bags and we keep getting more and more comments and all of these are you know thoughts and uh improvements is not really they go a little bit deeper and benign you know they tackle benign issues. Um so taking the example of the [03:35:53] adoption rates one may say well you know when historians look at the wave historical events we select those diffuse slowly and become important and when we think about the the what we [03:36:06] Wikipedia categorize as more uh impactful events in more recently maybe it's the faster one that got picked there. So things like that just as example uh to what extent do you guys think we can close our eyes and rest [03:36:20] easily and to what extent we still need lots of brown bags the seminars like this. Thank you. [03:36:25] >> Um well I'm happy we're doing this seminar. So uh let's keep doing that. [03:36:29] But I I think that um the uh uh I think you're it's a great question because I think that as we try to treat it in the paper is thinking of this as a measurement tool like a lot of other types of measurements like a survey or a ruler or whatever else you might use in [03:36:43] a scientific procedure which still means you need to have scientific discipline around it. Um and obviously especially because it's a newer method you want to have a particular rigor around that. So I think that like the like for example with the tech adoption stuff there's all sorts of still decisions that we have to make. maybe GPT will be able to make [03:36:58] them later on um as a research designer but we have to make about oh when do we want to cut off the data because we're afraid of skewing of that sort of adoption rate thing um what do we consider a historical significant technology what other data like um other data sets do we want to validate against [03:37:13] so all those decisions I think still need to make be made before and after I do think one shift um alluding to the um to the discussant um point is like I think that um that there is a uh a need [03:37:27] for understanding what we do with an explosion of data. Um I think that you know it just is very easy to measure a ton of attributes. So having some more um rigor and thoughtfulness about um what we measure or do we measure on like [03:37:40] a one sample and then only run it once we've done all our experimental methods on a leftout sample or things like that to to avoid like p hacking or other sorts of concerns that can that can emerge from just having so much data. [03:37:57] All right. Uh, I appreciate all these presentations. I have like a thousand questions, but I will limit myself to one. Uh, and this is for the authors of the Gabriel paper. So, um, I think the scope of the paper is very impressive, but I was kind of surprised about the [03:38:10] complete absence of references to the existing tools in NLP. In fact, like natural language processing doesn't appear in the paper at all. Um, and you know, like since 2018, we've had tools that do this with just encoder only [03:38:24] models. Um, and I was wondering if you've compared your data and accuracy to any BERT models either offtheshelf or fine-tune because you know it seems to me the gold standard in NLP is using uh [03:38:38] birectional encoders and just like that's my question and let me just say as a comment like you know these these these BERT models just run on our laptops. They don't require a server farm to to run and they basically do the same thing. And so seems to me like [03:38:52] maybe the future isn't just using these massive behemoths that we have to prompt to do the work, but to like build better small tools, small birds that that do the job and that we can just run on our computers. [03:39:07] >> Um, so um I think like in terms of comparisons like yeah, there's there's a few comparisons we do in the paper, but I think there's like more to be done. I do think part of the idea we're trying to get at here is something pretty orthogonal to what you can do with something like BERT, which is BERT or [03:39:22] some of these more recent um usages in um NLP are stuff you can train that you have to like have technical skill to train to match human labels or match human performance on one specific labeling task or a narrow range of labeling tasks. What we're trying to [03:39:36] accomplish with GBT um law LLM scale models is the general nature of human comprehension, the general idea of human labeling that you just ask the question, you don't train anything, you don't do any steps like that. And trying to [03:39:50] establish that as a competent method that that isn't to exclude that like a set of other methods might not be useful. Like for example, there's still cases where you want something really lightweight to analyze like 10 um 100 million documents and you need [03:40:03] something more lightweight than an LLM. But what we're trying to show is that LLMs are broadly capable at this sort of comprehension types task. Um and they do this labeling task um about as well as we would get from human labelers. Um [03:40:17] which is also the gold standard for BERT and other sorts of methods. So we try to show that generalizability, the technical ease for more researchers and accessibility. And then finally, um that it works for pretty small models that are still LLM grade. So a small LLMs can [03:40:32] still do this at scale. >> Okay, we're about out of time, but one quick question and quick answer. [03:40:37] >> Yeah. So for for Ben, so you showed us the statistic that uh your your more fancy model is doing better on average than the the ontheshelf model, but maybe I misunderstood it, but I didn't see [03:40:50] like how good is the model doing like in an absolute sense. [03:40:56] >> So it's it's actually a little hard to quantify. So the games are incredibly variable. Some of them are like the 511 game and some of them are the 2045 game. [03:41:05] And so like kind of doing traditional distance metrics are not like really interpretable over the vast majority of the games. Um the we have like a bunch of statistics about the optimized agents, you know, putting the majority of the probability on the choice the [03:41:20] people picked most like 80% of the time. And just in general, they're putting substantially more probability on the choices the human took most in the baseline. But honestly, it's a little bit of an open question about how to like give a really interpretable metric of how to compare so many games at once [03:41:34] with vastly different action spaces. We're still working on it. I'd love to talk about it. [03:41:39] >> Okay. Thanks. So, uh before we go to lunch, we're going to do a photo out front so future AIs will know we were all here at this moment. Um so, if you could please uh Elsa, is it right out there in in front there?