Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # General Social Agents Authors: Benjamin S. Manning, John J. Horton Discussant: Matthew O. Jackson Video: https://www.youtube.com/watch?v=VdvT0JzwMHU&t=7858s ## Talk (02:10:58 – 02:31:30) [02:10:58] >> All right, folks. Let's get uh started again. Um, so we're going to have uh Ben Manning uh presenting the first paper as before uh 20 minutes and uh write down your questions. We'll have time for questions at the end. Okay, take it [02:11:11] away. >> All right. Is this microphone on? Okay. [02:11:26] Hi everyone. My name is Ben Manning. I'm a PhD candidate at MIT. I'm really honored to have the privilege to present to you all today. I'm going to be sharing my paper general social agents. [02:11:35] This is joint work with my brilliant adviser John Horton. Broadly speaking, this paper is about using large language models and generative AI to kind of simulate human responses to surveys and experiments. I think it's a nice sequent [02:11:50] where we thought about a lot of survey data and classifying lots of tasks. So, for those of you totally unfamiliar with the idea of using an LLM to simulate human responses, I've got a little toy example for you here that I wrote up last night. The [02:12:03] idea is pretty simple. I'm sure some of you have done a little play acting with chat GPT or Claude before. One thing you can do with a large language model is open it up and prompt it as if it is some sort of human. This is myself. You can give it a bunch of little traits or characteristics and then ask it kind of [02:12:18] any social scientific question you might be interested in. I've ripped this off from a paper in the 80s by Danny Conorman and Richard Thaylor that looked at price gouging. So that's the basic idea, but it's pretty general. One does not need to be limited to kind of simple personas. One can endow these language [02:12:33] models with sort of any sort of observable traits, states of the world, or unobservable mental processes that one might want the language model to inhabit when it's making some sort of prediction about a setting. And this is exciting because these language models [02:12:48] and these simulations offer a unique combination of features that social scientists I believe should be really attracted to. And that is one they're extremely flexible. So if I want to make a prediction for some setting, a language model can do so for any setting that I can describe in natural language. [02:13:03] And then since it's a language model, it's really cheap and it's really fast. [02:13:07] One can imagine using this to kind of explore ideas, you know, augment survey data from some of the uh presentations we saw this morning, do AB tests, understand people's behavior, lots of different things. And I should note, you don't have to trust me that this is [02:13:20] potentially interesting. There are now dozens of papers demonstrating that these AI simulations can be highly accurate predictors of human responses. [02:13:28] This is a small subset of ones that I picked out. I'll also note that there's a lot of startups looking into this. In fact, there was one that came out of Stanford recently that just raised $und00 million and was for better or for worse talked about in the Wall Street Journal and New York Times a lot recently. [02:13:42] However, for every one of these papers demonstrating these AI simulations can be highly accurate predictors of human responses, there is another paper showing that the exact same LLM is a very poor predictor of human responses. [02:13:55] It'll caricature human responses. it will respond with inappropriate variability and sometimes it will hallucinate, for lack of a better phrase, give a ridiculous answer. Now, I kind of started off the presentation saying I thought that this tool was exciting and interesting, but I'll be at [02:14:10] first to admit that if we have inconsistency with a tool that sometimes works well and sometimes does not, that's a little bit of a knife in the heart for its generalizability and utility for us as economists in social sciences. So, as we're thinking about using this tool, there's a key challenge [02:14:24] here, which is reliably building agents and simulations that generalize. And that's what I'm going to briefly talk to you for about 17 minutes for the rest of today. I'm going to ask the research question, can we think about designing methods and general approaches to creating agents that will more [02:14:38] accurately predict human responses across a wide spectrum of settings? [02:14:43] Hopefully, by the end of this presentation, I'm going to convince you that we've started moving in the right direction on this. I'm also going to pull put a little bit of structure on this question because it is very broad. [02:14:52] Today I'm going to be focusing on prompts, not fine-tuning. I'm going to be thinking about matching distributions of human responses, not individual humans. And I'm going to be thinking about looking at settings where no prior human data exists. [02:15:06] Here's an overview. On the next slide, I'm going to present to you the approach, the general method we have come up with to improving these simulations. It's a little bit abstract. [02:15:14] So I found the best way to do this presentation is to walk you through the approach and then effectively use the rest of the presentation as a case study example of applying the approach. So if it seems a little little bit abstract, I'll get to it in a few minutes. I'm then going to introduce to you some behavioral economics games from the [02:15:29] literature. These are classic games from a famous AERT paper in from 2012. And I'm going to use these games. I'm going to create a very large set derived from a simple game as a playground to test my approach. I'm then going to show you some experiments where we test the [02:15:43] approach on human subjects, never before seen experiments. Then I'm going to talk about other applications and then I'm going to conclude. With that said, I'm going to get right into it. So suppose you want to generate those simulated predictions for some setting where you [02:15:58] have no prior human data. My claim here is that you should follow the following four steps. The first step in the approach is to identify some core theory that you believe explains or motivates behavior in your novel setting of interest. Although it's a novel setting, [02:16:13] presumably you have some sort of domain expertise, some sort of priors that you believe might dictate behavior in that setting. And you're going to create a bunch of those little personas like the one I showed you on the first slide. [02:16:23] Hopefully they don't have my name in it though, like the first one did. The next thing you're going to do is identify multiple and existent human data sets with different response distributions that you believe are also generally [02:16:35] dictated by the same underlying theory. Third, you're going to use some of those data sets, but not all of them to curate your original sample of agents. [02:16:47] I'm going to do an optimization procedure. I'll talk you through it in a few slides. And then finally, you'll use the remaining data to validate that curated set of agents. a sort of agentic cross validation. And the claim in this paper is that by following this approach, we can dramatically improve [02:17:01] the predictive power of simulations. All right, I'm going to introduce to you now a reasoning experiment. I'm going to introduce a game and then we're going to talk through applying the approach in this setting. I'm going to use a game from this AER paper, the 1120 money request game. This game is pretty [02:17:16] simple. It's also my favorite game in behavioral economics. Here are the instructions. There's one slide you pay attention to today. Please be this one. [02:17:23] The rest of the presentation will be hard to follow if you don't understand the game. The game's simple. Two players simultaneously choose an integer between 11 and 20. Players are then given the amount of money for the integer they choose. So, for example, if I pick 17 [02:17:38] and Professor Bulfson picks 19, I get $17, he gets $19. There's a catch. [02:17:44] There's a bonus opportunity. If I undercut my opponent by exactly one, I get plus 20. So if I pick 17 and professor Bernolson picks 18, I get 37, he gets 18. You can see there's this pressure to go higher to get guaranteed [02:17:58] money, but a strategic incentive to try to undercut. Now this game has a general structure that we can create a space of strategic games. You'll notice there's nothing special about 11 and 20 being the bounds [02:18:12] for this game. What could kind of choose anything for those bounds. One could also choose anything for the guaranteed amount of money or the bonus. Maybe instead of the 1120 game, we have the 514 game. Pick a number between 5 and [02:18:25] 14. You get a guaranteed amount of money plus two. I've kind of put in random rules. The 2237 game. I could go on and on and on. And so what I'm going to do, and I did in the paper, is created a space of these games, 800,000 of them. [02:18:40] And I also dramatically varied the bonus rule to make these strategic incentives different. I don't have time to walk you through each of these different bonus rules, but you can imagine creating a cartisian product of tons of games. And these games are where I'm going to want to make predictions. Besides some of the [02:18:54] games studied in a rotten Rubenstein, the vast majority of them have never been studied by humans. All right. So, if we want to make predictions in these games, a nice first step might be to just see what happens when we do have an AI simulation play the game from a Roden [02:19:08] Rubenstein where they've already studied an experiment. [02:19:12] What I did is I took the rules from a Roden Rubenstein. I gave them to GPT40 a thousand times with the temperature set to one and asked it to play this game. I used the real instructions from the paper. This is just kind of verb uh non-verbose instructions so you can get [02:19:26] the gist. I also added nothing else to the prompt. I just asked GPT40 to pick a number. And I'll note I'm going to call this the baseline or off-the-shelf LLM. [02:19:35] It's going to be my reference. Here's how GPT40 responded. It's an empirical distribution of its responses between 11 and 20. You can see it almost exclusively selected 19 across a lot of these runs. [02:19:48] Now, how do the humans respond in a rod and Rubenstein? They asked 108 students from the Hebrew University at Jerusalem to play this game. And here were their responses. The black bars are their empirical distribution that I've superimposed on top of the AI. As you [02:20:01] can see, they predominantly pick 17 and 18 with a little bit on 19. The takeaway is not like the LLM. [02:20:09] We're not feeling great about making predictions on those 800,000 new games if we can't even get predictions on the places where we have human data on something that is surely in the LLM's training corpus. So, what are we going to do? We're going to go to the approach. The approach has four steps. [02:20:23] We're going to apply them each in turn. First step, identify a core theory where we think plausibly dictates behavior in the novel games. [02:20:32] The model I'm going to use to generate that theory is level K reasoning. It's what a Rod and Rubenstein originally wanted to study in their paper and it is a theory that posits that people reason according to different levels in a population. There's lots of permutations. This is the one I'm talking about. I'll walk you through. A [02:20:46] level zero reasoner, for example, doesn't think about its opponent. They just pick the obvious number. We're playing the 1120 game. 20 gives me the highest amount of money. I'll pick it. A level one reasoner, my opponent is level zero, so they'll pick 20. How do I get [02:21:01] the best? I should pick 19. it gives me the $20 bonus. A level two reasoner thinks that they're playing a level one reasoner who knows they're playing a level zero reasoner. Keep going down. [02:21:11] You can see there's this cyclical process. The model more or less defines that a level K player best responds with the pure strategy according to playing a level K minus one player. So according to my approach, I just came up with 10 prompts that I think are broadly [02:21:25] representative of this theory. There's the 10 prompts. [02:21:31] The next step, find multiple distinct human data sets where we think the theory applies. Arod and Rubenstein, as I said, had humans play the basic version of the game, but they also had people play two other versions of the game. Fortunately for me, these are [02:21:45] relatively similar to the basic version. They vary the bonus rule a little bit and the guaranteed number of points. The big takeaway is that humans actually respond kind of differently to these three games. So, they offer kind of distinct tests if we have our agents potentially generalizing. [02:21:59] The next step in the approach is to optimize the original theoretically set of agents to match some of the human data from these games. And then the fourth step will validate in the others. [02:22:08] When I say optimized, what I'm going to do is I'm going to take each of these 10 prompts. You can see them on the page. [02:22:14] I'm going to give them to GPT40 independently, have each of those 10 prompts play the game a thousand times, and then I'm going to find the mixture of those response distributions that they generate that best approximates the original human data and a Roden Rubenstein. There are a lot of ways to [02:22:29] solve this optimization problem. I used some linear programming. [02:22:33] So once again prompted the model to create a response distribution for each game. Learned the weights that best optimized to the original human data. We get these weights. Happy to talk about them more later. Digging into them would take a lot of different time. And here [02:22:47] is the response distribution implied by solving this optimization problem. The black bars in both histograms are the same. It's the human data. And we can see the red is GPT4 out of the box and the blue is doing this optimization procedure curating a sample of agents. [02:23:01] As we can see, it's far closer. Of course, we were specifically solving an optimization procedure to bring the distribution towards the human subjects. [02:23:10] It's like in sample training error for machine learning. So the first step in the approach is to validate these agents. I'm going to take that exact same set of agents without any changes, that weighted mixture of them that I identified, and have them play the other two versions of the game that Arod and [02:23:25] Rubenstein have human subjects data for. Here is the basic version on top. I'm keeping the results there for reference. [02:23:34] We have the human subjects data is going to be consistent in every single row, black bars. The baseline is in red and the optimized is in blue for all of these histograms. [02:23:43] When I have GPT40 play the costless and cycle, those other two versions of the game from the paper, here's how they respond compared to the human subjects. Two things to take away from here. First of all, once again, the model always seems to pick 19 for the [02:23:57] most part. And you can also see that the modes have shifted substantially for the human subject data. Now, we're going to take the optimized agents optimized only to play the basic version of the game and see how they play the cycle and costless and how that matches the human [02:24:10] data. Here are the results. The blue histograms superimposed on the black human data consistent across rows. As you can see, the optimized agents are far closer. The costless is almost perfect. And these agents shift [02:24:25] appropriately despite the fact that there's a mode shift in both the AI response distribution for the optimized agents and the human subjects. For the basic version of the game, the mode is 17 for humans. For the costless version, the mode is 19. [02:24:40] Two key points here. They were optimized only to match the basic version of the game. Yet they generalized to the other two which means the agents are kind of by definition flexibly predicting responses across various sets of games. [02:24:52] How did we get agents like this? Well, we followed the four steps in the approach. We identified a core set of theories, the level K reasoning that we thought might dictate behavior in some of these settings. We identified three [02:25:05] existent but distinct human data sets. We optimized on some of those data sets and we validated on other sets. The central claim of the paper, as I said before, is constructing agents in this way can dramatically improve their predictive power in novel settings. I [02:25:20] have yet to predict anything that's novel. So, the next thing we did was put that to the test. We want to see how our agents do in a novel setting. [02:25:29] Remember those 800,000 games I introduced to you before? That's where I'm going to try and make predictions. [02:25:34] I'm going to try and predict initial human play in those 800,000 games. [02:25:38] Technically, the three games from Arod and Rubenstein are a subset here. I'm going to remove them. So, it's maybe a little less than 800,000. [02:25:46] We ran a human subjects experiment. I randomly sampled 1500 of those 800,000 games. I had GPD4 out of the box play the games 100 times. I pre-registered it. I then had the optimized sample only optimized on the original Arod and [02:26:00] Rubenstein data. play each of the games a 100 times pre-registered in it. And then we asked around four and a half thousand prolific workers to each play one of those dramatically varying 1500 games. We then paid out 1% of them. And we're going to see how well we predict [02:26:14] them. I should note that this incentive was large. Some people played the like 3545 game and then won and so they won like $55. So the bonuses were substantial if you hit. The way I'm going to evaluate these predictions is [02:26:28] using likelihood ratios. So my agent distributions are effectively a model for each game. We have a response distribution generated by the optimized AI and the baseline AI for each game. So I can calculate the joint likelihood that the human data for any given game [02:26:42] comes from the optimized or the baseline. I can then set this up as likelihood ratios. Two competing models, the log of the optimized over the baseline. And this will give me a distribution of per game likelihood ratios across all of these games. The [02:26:55] relative predictive power. Note that if a likelihood ratio is greater than zero, it means the optimized is better when I show you these results on the next slide. [02:27:04] So here I'm going to show you a histogram of the per game likelihood ratios. Zero means the optimized is better. I put them in green. Less than zero, the optimized is worse. Here is the histogram of the results. This is [02:27:18] the 1500 games. Their log likelihood ratios. The dash black line is the average. The standard errors are there. [02:27:24] They're just incredibly small. The interpretation here is that the optimized AI is a far better predictor of the human responses. It puts something like 50% more probability on the choices the human made than the [02:27:37] baseline GPT4 out of the box does. I'll also note something cool about the setup is that since the 1500 games were randomly sampled from the 1800 800,000 games in the population and we randomly assign some people to those games, we [02:27:52] get some nice econometric guarantees saying that this is actually valid over the entire population of games. We show this in the paper. Another thing I don't have time to talk to you about today that we do is we compare these optimized predictions to a variety of maybe [02:28:06] stronger benchmarks. Maybe you're not impressed that we just beat the LLM out of the box. We use the best theoretical models that we can come up with that can also generate predictions over such a wide variety of games. There's actually not that many models that can do this. [02:28:20] We use the cognitive hierarchy model from uh camera at all in 2004 and try various different uh values for the pan distribution if you're familiar. [02:28:32] As I end the presentation, I'd reiterate to you that this is the approach that I proposed that can dramatically improve the predictive power of AI simulations. [02:28:40] And we demonstrated that it worked this setting. In the paper, we also demonstrate trying to be good scientists what happens if we explicitly violate portions of the approach how everything goes. So we don't ground our original agents in theory and we show agents that [02:28:55] fail validation largely do not improve predictive power. [02:28:59] We also test this approach in an entirely different domain. Uh games derived from social preferences. These are drawn from a paper by Charnis and Raven in the early 2000s. We do the exact same thing. Go through optimize and validate with different theories on [02:29:13] subsets of games from their paper. We then make up a bunch of new games, pre-register our results, and see how we predict humans. I also have another setting that I've been working on as a part of the revision where we've been running opinion polling for consumer [02:29:26] views on sentiment relating to AI governments and advertising and we're asking people lots of different questions coming up with theories that we think motivate their answers optimizing and validating on some set subsets of questions and then seeing if we can predict the others. I think this [02:29:41] might be particularly particularly relevant for some of the things we saw in the second presentation this morning predicting people's views around AI. [02:29:49] I'll conclude with two slides. First of all, as I said in the two initial lies the presentation, AI simulations are powerful because they are extremely flexible. [02:30:09] data sets can really help. We're hopefully moving in the right direction. [02:30:13] We haven't solved everything. There's lots of open questions. Here's a few of the many questions I'm hoping to pursue. [02:30:20] First of all, I showed you one setting in depth where we think this works. I then briefly referenced two. We're hoping to scale it to hundreds if not thousands. I think that's what it will take to make this really compelling at scale. I also think it'll be interesting to think about how you can get better [02:30:34] measures of when things might generalize. We're focusing here on settings where we have no prior human data, where it can be hard to get guarantees of generalization without strong assumptions. But maybe there's new tools we could make through discussions at this conference to [02:30:48] understand when we might generalize better. And finally, I think an interesting question that I've yet to answer myself, but hopefully some of the mech interpretability, mechanistic interpretability researchers will, is thinking about what does it mean when a set of theoretically motivated prompts [02:31:03] improve predictive power. So, as I showed you today, using those level K prompts, I was able to improve the simulation's power. But since the LLM is a black box, it's hard to know how those actually went through. So, could we potentially search for more hypotheses [02:31:17] based on improving predictive power is a little bit of an open question. That's all I have for you today. Thank you so much for this opportunity. I really appreciate it. [02:31:26] [applause] All right. And our Thank you, Benjamin. ## Discussant remarks (03:13:13 – 03:28:17) *Shared across the three papers in this session; Matthew O. Jackson discussed all three.* [03:13:13] Great. So thanks to Eric and Karen for organizing this and to assigning me to three very exciting papers to discuss. [03:13:22] So I want to start just with this picture and to I I don't know how many of you know historians but historians a lot of what they trained themselves on was collecting data going through [03:13:35] archives spending time making sure they could translate qualitative things or undigitized things into things that could be eventually used and curating it and making sure that they could somehow validate it and then going on to to [03:13:50] using it to to do some studies. And so one question is, you know, how long did it take to do this compared to how long did it take historians to to spend time collecting this kind of information? It couldn't be done at the scale that was [03:14:04] was done here in in a matter of probably uh days or weeks um to put this together. So it's very impressive and the scale is eye opening. So AI aided [03:14:17] research is going to be lead to an explosion in the quantity um and the potential for quality of research. And so um we need ways to harness it and evaluate it. And I think that that's you [03:14:30] know something that comes through in all of these is making sure that we're understanding what we're getting out of it. And so I want to talk um in a little bit of detail about this. And I'll start with just a picture uh from a paper on AI behavioral science that puts this in [03:14:45] context and a number of the authors on this are are in the room. Um so this is from a a workshop last year and it sort of breaks things down into three different categories. So we can think of assessing AI behavior. So AI is out [03:14:59] there. It's it's going to be being used in increasingly um complex ways. We want to understand how it's behaved. That's sort of the alignment problem. Um, we're going to be using AI for behavioral science. So, uh, using it to to actually [03:15:13] understand behavior and to to study humans. And then there's also understanding human AI interactions. And that was more what we saw in the first session today. And so this these three [03:15:27] papers fall mainly in this category of of using AI for behavioral sciences. [03:15:33] And you know when we think about why AI is useful, it's useful to sort of think about what it expands in terms of why is it better than humans at different kinds of tasks. And I think in in this in these papers, what we see is it can [03:15:47] digest a lot more data. It can make predictions. It's going to be useful in simulations and we can design and control it. So these are sort of different ways in which AI is enhancing human behavior and it's useful to keep those in mind when we're trying to [03:16:01] evaluate this. Okay. So so AI enhanced research um it can enhance the simulation techniques we have. It can help us develop new modeling tools. It's going to allow us [03:16:14] to do larger, more complex, better trained and calibrated simulations. It's going to be much cheaper than humans. Um it can run many more experiments than we can. It can infer things from prompts [03:16:26] and uh try and we we can use that to understand behavior and I'll talk more about that. And then it also has this possibility for new data analytics that we saw in the last paper where it can quantify qualitative data. It can [03:16:40] digitize things. We have ways of of doing visualizations and categorizations that we didn't have before. So there's a whole series of ways in which it's going to be enhancing. And that's not to mention the way that we use it every day, which is just, you know, doing if you're a theorist, you do proofs and [03:16:55] conjectures or coding. You can do lit reviews. It can do a whole series of things. So, it's it's going to be coming into our lives increasingly as scientists, researchers, people developing engineers, etc. Um, the [03:17:08] papers that we saw in this session, I think it's useful to sort of take the high level and think about where do they fit? They fit in these categories. We're going to be building larger, more complex, better trained simulators and [03:17:21] models of behavior ways in which we can um emulate humans or imitate humans, I should say. Um we can infer from the prompts, sorry, we can infer from the prompts and the settings about these [03:17:34] induced behaviors. So, uh I I'll talk about that especially with with respect to to Ben's paper. There'll be things we'll be able to infer about theories from this. And then we can also quantify qualitative data and use it in ways that [03:17:48] we were not able to before. And so what do each of these papers teach us? So with three papers, it's hard to sort of spend a lot of time on each paper. So I'm going to try to try and pull out big points from each of these papers and think about what are we learning more [03:18:02] generally. And one thing I want to emphasize is that, you know, LLMs are temporary. So we we're we're we're looking at this sort of particular object that's come along in the in the last couple of years. Five years from now, we'll probably be talking about a [03:18:16] different form of of AI or an enhanced version which is going to evolve in in ways that we can't anticipate at this point in time. So I want to keep this at a level of you know what do we learn in general rather than what do we learn specifically about these instantiations [03:18:31] of AI that are are you know quickly um becoming ubiquitous. So um Ben's paper made clear that we can begin to harness the steerability of LLMs and more [03:18:44] generally steerability of any kind of system and we can train it in some settings and then see how it performs in other settings and that's going to be a general you know it used to be agent-based modeling now it's going to [03:18:57] be um LLM based modeling or or other kinds of models. The advantage of that is that it it allows us to see exactly what we're putting into the models in ways that that we can build prompts that [03:19:11] generalize and try and understand what we're inferring from that. And I want to point one point out one thing that I think is important in terms of of the of of Ben and John's paper is that when you think about the the example that they [03:19:26] did with the 1120 game, they were using level K reasoning. So level K reasoning allows you to produce a human behavior map that's very nice. Now you could have used some other model instead of you [03:19:39] know that comes out. So suppose instead we want to test whether humans are risk averse or whether they're fair or altruistic. You could have instead tried to put that into that model and see whether it would have done it well in terms of this you know predicting the [03:19:53] 1120 game and it probably would have failed miserably. And so so that allows us to say actually there's something about the depth of strategic reasoning that is present in these kinds of settings. And so now this is a tool for [03:20:07] saying you know if you want to understand how people are behaving in certain kinds of strategic settings where they're interacting with other people and trying to anticipate them you need a model like this. So it allows us to to sort of you know go through and [03:20:21] decipher as I put in the bottom decipher human behaviors and as you change the game you change the models you're using we can learn a lot about that. So I think that it's can be a very powerful [03:20:32] tool for not only simulating but also you know going back and figuring out how do we understand human behavior by what we had to do to train the LLM in order to you know give a certain kind of [03:20:46] behavior out. So it can be a very useful tool in that manner. [03:20:51] Um, and I I think you know that gets back to the to the lab rats that that um Abishek mentioned and and effectively we'll be building different kinds of lab rats and we actually learn a lot about human disease by figuring out what kinds [03:21:06] of lab rats did we have to build in order to test a medicine and so so there's a a nice you know feedback there and and in terms of um Abishek's talk I think you know again this is teaching us [03:21:20] something about the harnessability and the steerability of LLMs. It goes a little more into the to the black box of LLMs and I want to come back to that point in a minute. But using these [03:21:32] sparse autoenccorders, probes prompts to steer in increasingly interpretable ways. And and that interpretation word was mentioned a bunch of times. And I I'll say a few things about interpretability. It's going to be [03:21:46] important in terms of us understanding how these things are working. And as AI becomes increasingly complex and opaque, it's more and more important for us to be able to understand why is it doing [03:22:01] what it's doing and how do we interpret it. And um it's useful to you know the the the idea of working on llama rather than other LLMs allows you to work with with the um sparse autoenccoders and the [03:22:14] probes in ways that you can't with other AI and and having access to the understanding of how these systems are working is going to become increasingly important. So, not just building better simulators, but understanding why [03:22:29] they're simulating in the ways that they are. And and that's going to be an issue. Um, in terms of Gabriel and the the the the paper um on on using this as a a systematic LLM tool to measure [03:22:44] qualitative data. Again, I think this points to both the enormous potential of us being able to replace a lot of wrote [03:22:56] um repeatable tasks of, you know, quantifying data, collecting it, analyzing it at at scale in ways that are going to be consistent. You know, if you employ four different RAS to go and do this and they each have a slightly [03:23:10] different encoding techniques and so forth, you might end up with data that's much worse than using a very simple um system that you can give explicit directions and you know it's following very specific rules in how it's coding [03:23:24] things. So it can do it with accuracy, it can do it directly as was mentioned and the cost replicability and so forth the scale is is unmatched. Uh I think one thing that that comes up in sort of [03:23:37] thinking about this is this approach could be used to start dealing with a problem that we've had for years which is understanding measurement error out there in data data sets that are existing and have been used quite a bit. [03:23:51] Um you know this this morning's talk about building everything on one pillar. [03:23:55] um understanding to what extent are we sort of using data which was collected by humans or constructed by humans in ways that might be somewhat idiosyncratic. So are there ways that [03:24:08] that we can use this kind of tool to better understand and collect parallel data sets um other kinds of data sets and begin to to understand measurement error better. So I think there's a lot of promise in this that goes beyond just [03:24:22] using it as a single tool for single studies but but using it in more general ways to understand measurement. [03:24:30] Okay, I want to say a few things that that I think are are sort of interesting big picture um issues that we face as AI enters into our world of of data and and so forth. And one is a lesson from a [03:24:44] data boom that came in the 1980s in finance and people you know there was suddenly minute-by-minut stock data that was available. People started analyzing it like crazy. We found the day of the week effects, the January effect, the [03:24:56] index effect, um large cap effect. There were all kinds of of particular anomalies that were coming up in sort of shiny new facts. And as we start deploying this um these tools, it's [03:25:09] possible that we can become uh you know just inundated with different interesting relationships that we couldn't have seen before. And I think it's that means that we have to start disciplining ourselves or we're just [03:25:22] going to be under a sea of of potential research questions. And so somehow managing to to understand what we should be asking, how we should we be asking it rather than just, you know, throwing it out at everything we can see is is going [03:25:36] to be hard and theory in the in in this age of AI can still be very useful in sort of guiding the deployment. What are the questions we should be asking? How do we assess um AI's use and and output? [03:25:49] Um and and how do we enrich the models in ways that are useful? [03:25:54] And I think there's um one kind of interesting paradox I've been thinking about quite a bit and and I think that this is going to become increasingly um important and and you might think of this as like the chess go paradox. So [03:26:08] it's been just about 30 years now since uh Deep Blue beat the a grandmaster. Um it's been about 10 years since uh a Go Alpha Go um beat one of the best Go players in the world. So, you know, [03:26:22] we're getting to a point where AI can do things in some ways better than humans, at least in very specific kinds of tasks. Then the question becomes at some point if it's doing things that we can't understand. So, it it generates kinds of [03:26:37] plays that look very peculiar and so forth. So, that teaches us things, but at the same time, it might be that we can, you know, it's it's doing things that are beyond our abilities to to [03:26:48] comprehend. Then how do we make use of that? Do we just trust it? So if this is, you know, if we get a simulator that does better at weather prediction, that's fine. Maybe we don't need to know exactly how why it's predicting exactly what time it's going to rain tomorrow [03:27:03] and how much it's going to rain in different parts of California. Um, that would be fine. Um, what if we start using it for predicting what's going to happen to three different potential rate decisions by the Fed and we say, "Okay, [03:27:17] here's it can give you really accurate predictions. We can't tell you why it's predicting that, but we can tell you it's doing that. So, these systems are going to be do able to do that. Um, do we want to understand why a certain rate would would lead to one thing in the [03:27:31] economy versus another? Um, and you know, eventually when trade policy, you know, understanding things like trade policy and what the implications of trade policy are, we can be using these tools to create new simulations, to [03:27:44] analyze larger and larger data sets, to build these systems. And these systems are going to become increasingly opaque. [03:27:50] And so the kinds of tools that we're talked about today to measure, understand, and again this this word mechanic or you know term mechanical interpretability, we need to be able to look under the hood eventually if we [03:28:04] want to understand the and explain why things are working. And so that I think this an interesting paradox that we're going to have that it's going to be doing some things we'll never be able to interpret and other things that we will and you know how do we handle that ## General Q&A (03:28:17 – 03:41:40) *Shared across the three papers in this session.* [03:28:17] trade-off. So thanks [applause] >> really interesting set of questions from from Matt and from the presenters. So why don't we line up over here for questions and uh while you're all doing [03:28:31] that I'll take the uh moderator's privilege to to ask the first first question but go ahead and line up there. [03:28:37] Um, so it's interesting all these papers thinking about where these things are seem incredibly powerful in many ways. [03:28:44] Where does the power come from? I was disappointed to learn that it's not magic after all, but it seems to come from the the knowledge that's in it. And I think there's some a bad aspect to that and a good aspect. One that I'd love to have people discuss a little bit more is the risk of contamination. I was [03:28:58] actually surprised at how bad the uh U u LLM did just out of the box on 1120 because the answers were right in the corpus and they could have just looked it up but but you know and sometimes they do that and you don't realize it. [03:29:11] Um, but also there's there's some potential for it to be really powerful. [03:29:15] And that also has an interesting thing is as as the technologies get more and more powerful, they're going to read all of your papers as well and learn the four steps to making a good uh simulator and know the relevant theories that it [03:29:29] should be drawing on, including some that we didn't think of in all likelihood. And then you sort of start wondering at at some point will the will the um out of the box version actually be so much better because it's it's drawing on everything and then more that that any of us can think about. So I'd [03:29:43] be curious how how you react to to those pros and cons. [03:29:48] And it's that's a question for any one of the the papers who presenters who like to Yeah. Go ahead. [03:29:56] >> Is this working? Um I think it's a great question. I'd first of all say like I'd love for my paper to be obsolete. I think that would be kind of cool. Uh I don't have a perfect answer, but I think there's a thing I've been thinking about like good [03:30:09] memorization and bad memorization for these foundation models for making predictions. And I think since it's an economist crowd, I think that good memor like bad memorization would be a student who is learning game theory for the [03:30:23] first time and memorizes the fact that defect effect is the pure strategy Nash that you should always play. like good memorization would be a student memorizing that a Nash equilibrium is when everyone is best responding to one [03:30:37] another. And I don't have a perfect answer of how I would differentiate between how to learn those two different types of memorization. I think they are in effect both memorization. Maybe like a cognitive scientist will argue with me that like one is a learning [03:30:51] memorization. I don't know how how to articulate those differences. But I think us being able to specify when models, us being able to understand those two different types of memorization better is going to be is going to be key. And I I've been thinking about it a lot. I don't have a [03:31:05] perfect answer, but I think that's the distinction. [snorts] I think that there is a difference between memorization and contamination [03:31:19] in the sense that just because a model knows something doesn't mean it actually is using it in making a decision or something like that. Um so like just because it knows of a game or knows a story um or even knows the outcome of something like we do some tests in our appendix of our paper on this like [03:31:33] doesn't mean it uses that when we're asking it do a measurement task. So I think you have to be like more careful if that like you know in a prediction task for example where that's um more likely to happen um on prior data. But um but I think that like part of what we [03:31:47] see in our paper is that like even though it could have known some of these labels or known the game or known stuff like that that um the models are you know kind of like we are capable of compartmentalizing to a certain extent. [03:31:57] We don't understand that full extent but but I think that's there. Um just to quickly add to that um in in some past work we've uh we've like actually looked for memorization when we know that the model was actually trained on the data and even then those rates of [03:32:12] memorization are maybe in the 10 to 20%. And part so I was I have I think one should be worried about memorization in a statistical way. It's kind of a case by case uh case by case question. But I think the the rubric that like just [03:32:27] because a model is trained on something is necessarily spitting things out verbatim from the text doesn't really hold partially because these models are trying to compress all the information that they're trained on. And so even if they're trained on a bunch of things, I [03:32:40] don't remember like chapter six of a book I read even though I read it in if I read that in third grade. So the models have kind of that same feature. [03:32:48] And so we have to start to get at a prompt by prompt or a case by case basis of what are they relying on when responding to a particular answer and that gets a bit bit more trickier. [03:32:58] >> Questions. >> Okay. So I have two questions for Benjamin. Um so first I wanted to know if you've tried any of the new models with your methodology like if you use GPT.5.4 four or 5.5 and if the baseline [03:33:12] improves when you use a better model and if the difference between your method and the baseline changes if the model improves and then the second question I was wondering on so you talked about the first moment on the distributions and that it it gets better I was wondering [03:33:26] on the second moment because we were talking before about kind of the the variance on the responses and how there's less variance for the interviews for example and I was wondering if the if the responses be if there's as much variance as uh there is for human [03:33:41] responses on on the AI responses or or does the variant shrink or uh if there's any comment on that. [03:33:49] >> So for the first question, I haven't tried more advanced models on our games specifically, but I have tried more advanced models doing the approach in other settings and I know other people have too. And I've I don't see a clear pattern that [03:34:03] the LLMs do a better job of predicting human responses unless human responses are like more rational. In that case, better models do tend to predict it. So, but in games like the 1120 game where it's hard to say what the most rational [03:34:18] thing is because the rational thing to do is to know how your opponent what your opponent is. Like I am rational to pick 18 if I know my opponent is a 19 picker. Um, so to that point, not not necessarily seeing it better across [03:34:32] settings and have tried the approach, although not for the specific games on other models. And then for the variance, uh, I haven't explicitly checked on the variance. I think you could kind of extend this to trying to match as many moments as possible as part of the procedure. I think that'd be like a cool [03:34:46] extension. And I think kind of the kind of what I would like to take what I would hope I would be conveying from the method is not that like one had to follow a particular optimization approach when trying to set up these uh simulations. I'm a little agnostic to [03:35:00] the optimization approach. I think there's lots of things you could try. [03:35:03] I'm guessing the more data you have, the more moments you could match probably means better, but it's an empirical question we'll have to explore. [03:35:11] >> Thanks. Okay, this is a question um about the application of a Gabriel. It's really interesting application we have here. My question is uh I want to take the author [03:35:23] stance on to what extent the post application human involvement are needed after you assemble the data. To take an example you know when we write papers there is a first stage um you know we circulate within the department in brown [03:35:38] bags and we keep getting more and more comments and all of these are you know thoughts and uh improvements is not really they go a little bit deeper and benign you know they tackle benign issues. Um so taking the example of the [03:35:53] adoption rates one may say well you know when historians look at the wave historical events we select those diffuse slowly and become important and when we think about the the what we [03:36:06] Wikipedia categorize as more uh impactful events in more recently maybe it's the faster one that got picked there. So things like that just as example uh to what extent do you guys think we can close our eyes and rest [03:36:20] easily and to what extent we still need lots of brown bags the seminars like this. Thank you. [03:36:25] >> Um well I'm happy we're doing this seminar. So uh let's keep doing that. [03:36:29] But I I think that um the uh uh I think you're it's a great question because I think that as we try to treat it in the paper is thinking of this as a measurement tool like a lot of other types of measurements like a survey or a ruler or whatever else you might use in [03:36:43] a scientific procedure which still means you need to have scientific discipline around it. Um and obviously especially because it's a newer method you want to have a particular rigor around that. So I think that like the like for example with the tech adoption stuff there's all sorts of still decisions that we have to make. maybe GPT will be able to make [03:36:58] them later on um as a research designer but we have to make about oh when do we want to cut off the data because we're afraid of skewing of that sort of adoption rate thing um what do we consider a historical significant technology what other data like um other data sets do we want to validate against [03:37:13] so all those decisions I think still need to make be made before and after I do think one shift um alluding to the um to the discussant um point is like I think that um that there is a uh a need [03:37:27] for understanding what we do with an explosion of data. Um I think that you know it just is very easy to measure a ton of attributes. So having some more um rigor and thoughtfulness about um what we measure or do we measure on like [03:37:40] a one sample and then only run it once we've done all our experimental methods on a leftout sample or things like that to to avoid like p hacking or other sorts of concerns that can that can emerge from just having so much data. [03:37:57] All right. Uh, I appreciate all these presentations. I have like a thousand questions, but I will limit myself to one. Uh, and this is for the authors of the Gabriel paper. So, um, I think the scope of the paper is very impressive, but I was kind of surprised about the [03:38:10] complete absence of references to the existing tools in NLP. In fact, like natural language processing doesn't appear in the paper at all. Um, and you know, like since 2018, we've had tools that do this with just encoder only [03:38:24] models. Um, and I was wondering if you've compared your data and accuracy to any BERT models either offtheshelf or fine-tune because you know it seems to me the gold standard in NLP is using uh [03:38:38] birectional encoders and just like that's my question and let me just say as a comment like you know these these these BERT models just run on our laptops. They don't require a server farm to to run and they basically do the same thing. And so seems to me like [03:38:52] maybe the future isn't just using these massive behemoths that we have to prompt to do the work, but to like build better small tools, small birds that that do the job and that we can just run on our computers. [03:39:07] >> Um, so um I think like in terms of comparisons like yeah, there's there's a few comparisons we do in the paper, but I think there's like more to be done. I do think part of the idea we're trying to get at here is something pretty orthogonal to what you can do with something like BERT, which is BERT or [03:39:22] some of these more recent um usages in um NLP are stuff you can train that you have to like have technical skill to train to match human labels or match human performance on one specific labeling task or a narrow range of labeling tasks. What we're trying to [03:39:36] accomplish with GBT um law LLM scale models is the general nature of human comprehension, the general idea of human labeling that you just ask the question, you don't train anything, you don't do any steps like that. And trying to [03:39:50] establish that as a competent method that that isn't to exclude that like a set of other methods might not be useful. Like for example, there's still cases where you want something really lightweight to analyze like 10 um 100 million documents and you need [03:40:03] something more lightweight than an LLM. But what we're trying to show is that LLMs are broadly capable at this sort of comprehension types task. Um and they do this labeling task um about as well as we would get from human labelers. Um [03:40:17] which is also the gold standard for BERT and other sorts of methods. So we try to show that generalizability, the technical ease for more researchers and accessibility. And then finally, um that it works for pretty small models that are still LLM grade. So a small LLMs can [03:40:32] still do this at scale. >> Okay, we're about out of time, but one quick question and quick answer. [03:40:37] >> Yeah. So for for Ben, so you showed us the statistic that uh your your more fancy model is doing better on average than the the ontheshelf model, but maybe I misunderstood it, but I didn't see [03:40:50] like how good is the model doing like in an absolute sense. [03:40:56] >> So it's it's actually a little hard to quantify. So the games are incredibly variable. Some of them are like the 511 game and some of them are the 2045 game. [03:41:05] And so like kind of doing traditional distance metrics are not like really interpretable over the vast majority of the games. Um the we have like a bunch of statistics about the optimized agents, you know, putting the majority of the probability on the choice the [03:41:20] people picked most like 80% of the time. And just in general, they're putting substantially more probability on the choices the human took most in the baseline. But honestly, it's a little bit of an open question about how to like give a really interpretable metric of how to compare so many games at once [03:41:34] with vastly different action spaces. We're still working on it. I'd love to talk about it. [03:41:39] >> Okay. Thanks. So, uh before we go to lunch, we're going to do a photo out front so future AIs will know we were all here at this moment. Um so, if you could please uh Elsa, is it right out there in in front there?