Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # Interpreting and Steering LLM Agents for Social Simulations Authors: Jiayue Fan, Arul Murugan, Shreyas Krishnan, Abhishek Nagaraj Discussant: Matthew O. Jackson Video: https://www.youtube.com/watch?v=VdvT0JzwMHU&t=9090s ## Talk (02:31:30 – 02:52:12) [02:31:31] And our next paper is interpreting and steering LLM agents for social simulations. [02:31:44] >> Okay, this works. I guess you wanted me to stand here. Okay. [02:31:48] >> Uh thank you for that, Ben. And and I think a really good uh segue uh to kind of what we're going to talk about is we're going to continue the theme of uh of social simulations with with large language models. Um our starting point [02:32:02] will be uh another old theoretical paper. This one uh this one that Oops. [02:32:11] I think this stopped working. Oops. Okay. Uh this one from Tom Shelling um um in the 60s and 70s where he was trying to help understand uh models of segregation. uh and and the what was most interesting to me about the [02:32:25] methodology was not um empirical data or even closed form economic models but really agent-based models simple rules applied to individual agents and the key idea is that these simple rules can then [02:32:38] lead to big aggregations such as explaining things like uh the separation between the the blue and the and the red dots for example in the in this in this period right and so there's this hope I think when if you can in indeed simulate [02:32:51] individual human preferences from large language models that you can actually start to get another way of theorizing that that used to be quite popular but more recently has gone out of style within uh within economics. And so what we've really seen is the applicability [02:33:05] of these general social simulations in uh for for a variety of different reasons that that Ben talked about. Uh and so I won't belabor them, but in particular this idea of scalability is particularly interesting to me is that there's lots of experiments, there's [02:33:19] lots of examples that we would want to test with human subjects, but the cost of doing so as well as the kinds of interventions that we want to run are often impossible with human subjects. So I see this as a new methodology that almost uh sits before it's it's a kind [02:33:34] of theorizing that that Shelling did that helps us generate new hypotheses, new understandings of humans um uh that we can then take to design much better or better targeted human interventions. [02:33:46] U and this includes not just the scalability aspect but also the type of hetrogenity that that Ben was talking about our discussion also has some great great work in this area of how do we generate synthetic hetrogenity with these LLM agents. so that we can better [02:34:00] understand human preferences and the heterogeneity of human preferences. And so we've seen that the my mental model for this is that neural nets kind of offer a lab rat for for social scientists. For a long time, our colleagues on the other side of the [02:34:14] university have had the ability to test their ideas and theories uh in in uh within these within these mouse models uh really cheaply and at scale before going out to the field. And so so the vision here is that kind of [02:34:27] complimentarity between methods that these kinds of large language models operate. And so if you take that analogy seriously um and you go look at kind of how that that research operates, biologists don't just use vanilla mice, [02:34:39] right? So they don't just uh work with uh mice that you find in the wild, but these mice are often engineered to help understand very specific questions. So, and I think Ben's paper was a really good uh starting point at that vision [02:34:54] where we don't just use something like GPT40 off the shelf in order to answer the questions that we want, but we want to engineer these agents in a way that are specifically geared towards building better social understanding. So, the 2007 Nobel Prize was given for these [02:35:09] transgenic mice, mice that are uh particularly modified, right? So in this case the the way that we're moving away from uh from uh offtheshelf mice if you would like is we're actually modifying with their internals uh with their [02:35:24] genetic structure um and then giving birth to mice that are specifically engineered for studying particular diseases. At this point you can go uh to any supplier and basically get a menu of mice in order to study any particular [02:35:37] question. So the vision is if I want if I have a theoretical model about risk preferences, can I go and get a risky LLM and order one from uh from from one of the from one of the companies. Uh unfortunately I don't think this is [02:35:50] going to be a very big business model for for the foundation lab companies. [02:35:54] And so something that we academics ourselves have to build is a menu of these social simulators so that we can better advance theory. And that's really the goal of this paper is to provide one vision of what that might look like. How [02:36:06] do we start going beyond uh of the shelf mice to maybe particularly engineered mice in a particular setting? And so the way I think about this is what we want to do is we want to open up the black box. So what happens inside of these models as they're responding? So as [02:36:20] they're as as they're giving me a preference, for example, in a particular situation that I prefer vanilla over chocolate, what might be going on inside the model that leads the model to choose vanilla or chocolate and what drives the distribution in those preferences over [02:36:34] time? And the way we frame it in the paper is we actually want two things from uh from these models. We want interpretability. So we want them to understand why is it that they chose one uh option over another. And second we want controllability. Right? That's what [02:36:49] these u these genetically modified mice do is not only uh do they help me understand why somebody picked vanilla or chocolate, but I should be able to dictate that a particular agent picks one over the other and it does that consistently. I think I think this we [02:37:03] refer to this idea as generalizability or predictability of these models. So not only one I want to study their predictability but I want to engineer that predictability. Um and so uh this uh Ben again referenced this term [02:37:16] mechanistic interpretability. It's a rapidly growing field uh in computer science and the basic idea is to look inside the model and I'll say a little bit more about what that means and use that for all kinds of practical computer science applications. So that includes [02:37:30] AI safety and alignment. For example, if a model is uh thinking about uh devising a nuclear bomb, uh I might want to deactivate that capability. I might want to understand why models fail in certain situations. Uh and this idea of steering [02:37:45] and control has also been examined, right? And so there are at this point a few different uh few different attempts uh in the field to use the internals of a model for all kinds of applications but none of those applications have [02:37:59] really had a social scientific uh topic or comport to them and that's precisely what we want to do in this study. So for those of you who are not familiar with this field and I I don't blame you it's pretty pretty recent uh perhaps its [02:38:11] poster child uh is the so-called golden gate claude. Uh so a couple of years ago, Enthropic put out this paper and blog post where they uh engineered a version of Claude that was completely in love with the Golden Gate Bridge. So [02:38:25] they said you could actually engineer the internal activations of the model and no matter what question you asked it, it would refer to the Golden Gate Bridge. And they showed that you can have that kind of predictable behavior inside of the model if you engineered it that particular way. In our example, uh [02:38:40] we're going to use uh models that go not just at objects, but at specific features of economic interest, such as things like risk aversion or trust. Kind of these theoretical ideas that we've [02:38:53] long written down in our models, but never been able to actually study. So here is a version where I can actually find capabilities or parameters inside the model that relate to these specific fundamental characteristics and then [02:39:07] actually turn them up or turn them down. And so the idea is that is this just science fiction or can this actually work in practice? And I'm going to show you a few examples where it can actually be a really promising approach. And so [02:39:19] so how can we engineer um these models so that we get interpretability and controllability when we do social simulations with LLM? That's going to be our focus. Uh just a quick reminder, we won't have a ton of time into the [02:39:32] internals of how LLMs work. But one of the key p pieces is a step four where there's this residual stream of activations, right? So as we put in an input such as a cat doing bajillion calculations in order to generate the [02:39:46] next token which in this case is Matt. And so the core insight of mechanistic interpretability is that those internal tokens as they move through the different layers of the transformer might be able to be read might have some interpretation in order for me to [02:40:00] understand whether the output is going to be a cat or a mat. Uh, and if I can do that, then maybe I can not just read out that preference, but I can maybe change the mat to a ball or something else in the future. Okay. And so we're [02:40:15] going to try So that's precisely what we're going to do is we're going to put these LLMs inside of social simulations and we're going to layer on these interpretability techniques. We're going to mess with the model internals. So we're going to both read out the model internals so we can understand what [02:40:29] these models are thinking and then we're going to uh pertur them in order to get predictability of options. I'm going to show you how that helps us really design better social simulations. In particular, we're going to test two different approaches from the computer [02:40:43] science literature. Uh the the more recent one that has been popular and that was a part of Golden Gate Cloud are these so-called SAEs or sparse autoenccoders. I'm going to talk about that in a second. But there's also an older technique uh called linear probes [02:40:56] which is much simpler. One thing you might want to uh kind of as a broad way to distinguish between these two approaches is think about approach one as kind of an unsupervised approach. So even before I'm going out and uh training uh or or even before I go out [02:41:10] and and and trying to use the model for any kind of applicability, I train something called a sparse autoenccoder. [02:41:16] And what that does is takes the model internals and maps it on to a set of features. The model that we use will have about 65,000 features. And that's kind of independent of no matter what game it is that you're using in or what [02:41:30] application you're using it for. So if you look at step four here, there are these internal features such as economic trade-offs, fairness, equity, rejection, so on so forth. The drawback of this approach is that this can be quite costly. It can say cost tens of [02:41:43] thousands of dollars to train. But once this is trained, it can be used across all of these parameters quite flexibly. [02:41:50] The other options kind of the classical approach are these linear probes uh where basically what you look at is you you train the probe for a particular feature that you care about. So in this case for example we have examples of [02:42:02] altruism. So we feed the model with examples where was choosing between being altruistic versus not and we get some natural variation in the model's responses. And so we train a simple classifier on top of that and say these [02:42:17] kinds of activations correspond to altruistic behavior and these other kinds of activations correspond to less altruistic behavior. And so I have a very specifically engineered probe just for altruism. So that's great but it doesn't work for anything else. So what [02:42:30] we're going to use is compare these two approaches to the more accessible way of prompting uh which is give give model certain traits. for example, tell the model that you are an altruistic individual or tell the model that you're a risk-loving individual. And we're [02:42:44] going to compare these kind of I would I guess like a more interpretability oriented approach with a more prompting approach. Um, okay. So, in particular, we're going to go over four different games. Um, I'm going to talk I think I'm going to have time mostly to talk about [02:42:58] uh lottery choice. uh but then I'll also talk about uh ultimatum games which we think about as games of preferences understanding whether models inherently are riskloving uh or risk averse uh whether they trust the other party or not but we also look at things like [02:43:12] capability where we want to understand how creative are models many people are using these models in order for IDI generation that that to me seems like inherently a different kind of task uh okay so so here's how that works so we [02:43:26] use uh uh John Horton's uh EDSL uh where we uh use that as a basic framework for for simulations uh and on top of that we have models take all kinds take basically go through lots of lotteryies [02:43:41] ultimatum games creativity games and product innovation on top of that uh the underlying base model here is llama so you'll notice that we aren't using uh some of the closed source models because you can't use them for these kinds of techniques you need to be able to look [02:43:54] inside of the blackbox to be able to apply them and so um all of the simulations that I'll be showing you will be based off of the Lama 3.370 billion model. Um on top of that, we'll train these two specific tools, the SAPE [02:44:08] um as well as the probe. Okay, so this is what the lottery game looks like. Uh again, important to understand. Uh we give the models a choice between a guaranteed 50 tokens versus a probability of some set of tokens. [02:44:22] Right? Right. So if a model is purely risk averse uh purely risk neutral is going at the level 100 is going to be indifferent between choosing the lottery or a guaranteed $50 anything above 100 it chooses the lottery. Anything below [02:44:35] 100 it chooses the safe option. But of course if the model is risk loving or risk averse the 100 number goes up or down. So the question is can we engineer models at a very particular level of risk aversion? Can I engineer a model [02:44:49] that at 120 prefers uh prefers the lottery but below 120 it prefers a safe option? Right? So then I can quantify the level of risk aversion of that particular model. And so this is what happens in our game with the baseline [02:45:03] model. So llama is a little bit risk averse. So it chooses to switch to the uh to the risky option at about 120 with prompting you. Uh so here we prompt the model to be you're slightly riskloving. [02:45:16] it switches completely to always preferring the lottery. So even though I guaranteed it uh even though I guaranteed it $50, 50 tokens, it switches to the lottery way before that in some ways, right? And so that's kind of not the behavior that one would want [02:45:30] if one was trying to engineer some degree of risk uh risk loving or risk risk aversion in the model. Um and so one of the cool things you can do with meurp is you can read out the internals of the model through these SAPE [02:45:43] features. So here are the top activated features as the model is thinking about the lottery. And what you can see is the most activated features have nothing to do with risk aversion. They're more about well looks like I'm faced with a situation with multiple choices. How do I decide between that? So that's a [02:45:58] feature inside of the model. But pretty high up there as features around calculated risk takingaking. So what I can do is engineer the SAPE to move that up. I can say well I want you to be particularly riskloving in this particular situation. And so we use this [02:46:13] pre-trained uh goodfire SAPE on top of llama 3.3. Uh and when we use that when we use that model, what we find is that these features go to the very top of what's activated. And even more [02:46:26] interestingly now that model behaves in kind of a more systematic way. It does take more risk, but somewhere in the middle. I can also engineer a different degree of risk aversion. And what you see is that with these pre-trained SAEs, [02:46:39] you get some variation but not precise control. Right? So I get some control uh in a way that I understand but I can't still engineer the model to have a exact level of risk aversion. For that you [02:46:51] need probes. And so for probes now I'm going to throw away the generalizability or the general utility of SAEs. Instead I'm going to say I'm going to engineer this kind of classification model just for this one task. And so I'll skip some of some of the details of this, but what [02:47:06] you can see here, the colors represent what the probe was trained for. That's the level of risk aversion that the probe was trained to uh to exhibit. And the curves that you see are the practical behavior of those uh of those [02:47:19] models. So now I basically can give you a menu of agents exhibiting whatever level of risk aversion that you want based on this kind of probe train, right? And so that's a level of predictability and control kind of this [02:47:32] engineered mouse that might be a very useful input into any kind of social social situation. We show the same thing for interpretability with the ultimatum game. Again, there's a feature for altruistic behaviors. Again, we show [02:47:46] that steering is helpful. Uh and these kinds of linear probes can add a lot of granular control. uh and so so the general methodology here that seems to be working for us is one where we use the SAEs to understand inductively the [02:48:01] inherent uh motivations and then use the probes for very precise precise control. [02:48:06] How much time do I have? Uh >> oh. Oh, perfect. Thank you. I appreciate it. Okay, great. Uh so um what happens when we go to other types of games? So uh so here we look at uh from psychology games of diversion creativity. So what's [02:48:21] commonly done in diversion creativity is people are brought into the lab and given a limited amount of time and say give me all the unusual uses for a brick or list the different ways in which you can use a stapler. Um and so what we do [02:48:34] is we again take llama and we we uh we have it go through all of these different tasks. Um and then we evaluate the so what's harder about this task is unlike the lottery there is no quantitative evaluation. So we have to [02:48:48] use kind of an LLM as a judge to score uh the creativity responses on a 0ero to 10. And so this is one of the other things we find in the setting is uh your ability to control will be a function of how well you can validate or check the model responses because you're training [02:49:02] the model for that kind of validation. Um and again with these tasks we can read out one thing interesting in llama is it thinks about the Russian word for brick as it's thinking about lots of different uses for brick. Uh but again I can change that. So with SAES I can make [02:49:16] the most activated feature the feature around professional innovation and creative problem solving. It turns out that that does not work so well with the creativity task. So SAEs are a pretty blunt instrument relative to prompting in changing how much more creative these [02:49:31] models get. So SAEs are again very helpful in understanding activations but much less helpful in getting predictability in future returns. But again here the probes do uh do really well. So we train the probe particularly [02:49:44] on being creative on this uh on this brick task and I have a target creativity score. So again I have more creative versus less creative LLMs for this brick task. And indeed the more creative LLM uh has many more uses for a [02:49:59] brick than the less creative LLM. And this also works kind of out of sample. [02:50:02] If you change the object or if you change the kind of application, the creative LLM continues to be uh much more creative in some of these some of these high creativity zones. Okay. So, so what did we learn and and take away from these sets of experiments? So, [02:50:17] first compared to prom personas interp methods give social scientists just a more direct way to open up the black box of social of social simulations. So I'm excited as we move. I think there was kind of this early excitement about simulations that this can can do [02:50:32] everything and then there's some kind of disillusionment that they in lots of cases don't really uh simulate what it is that we want to do and I think there's clearly a ton of value in them but we need to use them with care and developing better control methods [02:50:46] including uh the previous paper uh but also mechan offers another way in which we we can get kind of standardizability u in uh in in social simulations in particular I I think our other contribution was to introduce these two [02:51:00] specific tools. Uh Sapes help with interpretation, right? So they decompose activations into human readable features making them useful for inductive discovery. So in this case, I knew that lotteryies are about risk aversion, but maybe it's about two different kinds of [02:51:14] ads and I don't know why one ad works better than the other. Essays might help develop hypothesis about why that happens. Uh and essay steering works well especially for these kind of preferenceoriented tasks rather than the capability oriented tasks. [02:51:27] uh but um and linear probes on the other hand are completely useless for interpretation right so they need labeled data I need to know exanti what what it is what kind of behavior I'm looking for but if you can actually do that they can be very precise uh in [02:51:41] terms of the kinds of steering that you can get and so the practical recipe seems to be to use essays to discover model is using and use probes to steer what you can define so to conclude I'm excited about a future where we kind of might have a menu a digital kind of [02:51:55] transgenic mouse for social simulation ations and so that's an exciting possibility. Uh thank you to my co-authors Gavio who led a lot of the work and Aurel who's also here in the room as well as Shas. Thank you so much. [02:52:05] >> All right. Thank you. [applause] >> Third presentation will be uh GPTs as a measurement tool and then we'll have discussion and and questions. ## Discussant remarks (03:13:13 – 03:28:17) *Shared across the three papers in this session; Matthew O. Jackson discussed all three.* [03:13:13] Great. So thanks to Eric and Karen for organizing this and to assigning me to three very exciting papers to discuss. [03:13:22] So I want to start just with this picture and to I I don't know how many of you know historians but historians a lot of what they trained themselves on was collecting data going through [03:13:35] archives spending time making sure they could translate qualitative things or undigitized things into things that could be eventually used and curating it and making sure that they could somehow validate it and then going on to to [03:13:50] using it to to do some studies. And so one question is, you know, how long did it take to do this compared to how long did it take historians to to spend time collecting this kind of information? It couldn't be done at the scale that was [03:14:04] was done here in in a matter of probably uh days or weeks um to put this together. So it's very impressive and the scale is eye opening. So AI aided [03:14:17] research is going to be lead to an explosion in the quantity um and the potential for quality of research. And so um we need ways to harness it and evaluate it. And I think that that's you [03:14:30] know something that comes through in all of these is making sure that we're understanding what we're getting out of it. And so I want to talk um in a little bit of detail about this. And I'll start with just a picture uh from a paper on AI behavioral science that puts this in [03:14:45] context and a number of the authors on this are are in the room. Um so this is from a a workshop last year and it sort of breaks things down into three different categories. So we can think of assessing AI behavior. So AI is out [03:14:59] there. It's it's going to be being used in increasingly um complex ways. We want to understand how it's behaved. That's sort of the alignment problem. Um, we're going to be using AI for behavioral science. So, uh, using it to to actually [03:15:13] understand behavior and to to study humans. And then there's also understanding human AI interactions. And that was more what we saw in the first session today. And so this these three [03:15:27] papers fall mainly in this category of of using AI for behavioral sciences. [03:15:33] And you know when we think about why AI is useful, it's useful to sort of think about what it expands in terms of why is it better than humans at different kinds of tasks. And I think in in this in these papers, what we see is it can [03:15:47] digest a lot more data. It can make predictions. It's going to be useful in simulations and we can design and control it. So these are sort of different ways in which AI is enhancing human behavior and it's useful to keep those in mind when we're trying to [03:16:01] evaluate this. Okay. So so AI enhanced research um it can enhance the simulation techniques we have. It can help us develop new modeling tools. It's going to allow us [03:16:14] to do larger, more complex, better trained and calibrated simulations. It's going to be much cheaper than humans. Um it can run many more experiments than we can. It can infer things from prompts [03:16:26] and uh try and we we can use that to understand behavior and I'll talk more about that. And then it also has this possibility for new data analytics that we saw in the last paper where it can quantify qualitative data. It can [03:16:40] digitize things. We have ways of of doing visualizations and categorizations that we didn't have before. So there's a whole series of ways in which it's going to be enhancing. And that's not to mention the way that we use it every day, which is just, you know, doing if you're a theorist, you do proofs and [03:16:55] conjectures or coding. You can do lit reviews. It can do a whole series of things. So, it's it's going to be coming into our lives increasingly as scientists, researchers, people developing engineers, etc. Um, the [03:17:08] papers that we saw in this session, I think it's useful to sort of take the high level and think about where do they fit? They fit in these categories. We're going to be building larger, more complex, better trained simulators and [03:17:21] models of behavior ways in which we can um emulate humans or imitate humans, I should say. Um we can infer from the prompts, sorry, we can infer from the prompts and the settings about these [03:17:34] induced behaviors. So, uh I I'll talk about that especially with with respect to to Ben's paper. There'll be things we'll be able to infer about theories from this. And then we can also quantify qualitative data and use it in ways that [03:17:48] we were not able to before. And so what do each of these papers teach us? So with three papers, it's hard to sort of spend a lot of time on each paper. So I'm going to try to try and pull out big points from each of these papers and think about what are we learning more [03:18:02] generally. And one thing I want to emphasize is that, you know, LLMs are temporary. So we we're we're we're looking at this sort of particular object that's come along in the in the last couple of years. Five years from now, we'll probably be talking about a [03:18:16] different form of of AI or an enhanced version which is going to evolve in in ways that we can't anticipate at this point in time. So I want to keep this at a level of you know what do we learn in general rather than what do we learn specifically about these instantiations [03:18:31] of AI that are are you know quickly um becoming ubiquitous. So um Ben's paper made clear that we can begin to harness the steerability of LLMs and more [03:18:44] generally steerability of any kind of system and we can train it in some settings and then see how it performs in other settings and that's going to be a general you know it used to be agent-based modeling now it's going to [03:18:57] be um LLM based modeling or or other kinds of models. The advantage of that is that it it allows us to see exactly what we're putting into the models in ways that that we can build prompts that [03:19:11] generalize and try and understand what we're inferring from that. And I want to point one point out one thing that I think is important in terms of of the of of Ben and John's paper is that when you think about the the example that they [03:19:26] did with the 1120 game, they were using level K reasoning. So level K reasoning allows you to produce a human behavior map that's very nice. Now you could have used some other model instead of you [03:19:39] know that comes out. So suppose instead we want to test whether humans are risk averse or whether they're fair or altruistic. You could have instead tried to put that into that model and see whether it would have done it well in terms of this you know predicting the [03:19:53] 1120 game and it probably would have failed miserably. And so so that allows us to say actually there's something about the depth of strategic reasoning that is present in these kinds of settings. And so now this is a tool for [03:20:07] saying you know if you want to understand how people are behaving in certain kinds of strategic settings where they're interacting with other people and trying to anticipate them you need a model like this. So it allows us to to sort of you know go through and [03:20:21] decipher as I put in the bottom decipher human behaviors and as you change the game you change the models you're using we can learn a lot about that. So I think that it's can be a very powerful [03:20:32] tool for not only simulating but also you know going back and figuring out how do we understand human behavior by what we had to do to train the LLM in order to you know give a certain kind of [03:20:46] behavior out. So it can be a very useful tool in that manner. [03:20:51] Um, and I I think you know that gets back to the to the lab rats that that um Abishek mentioned and and effectively we'll be building different kinds of lab rats and we actually learn a lot about human disease by figuring out what kinds [03:21:06] of lab rats did we have to build in order to test a medicine and so so there's a a nice you know feedback there and and in terms of um Abishek's talk I think you know again this is teaching us [03:21:20] something about the harnessability and the steerability of LLMs. It goes a little more into the to the black box of LLMs and I want to come back to that point in a minute. But using these [03:21:32] sparse autoenccorders, probes prompts to steer in increasingly interpretable ways. And and that interpretation word was mentioned a bunch of times. And I I'll say a few things about interpretability. It's going to be [03:21:46] important in terms of us understanding how these things are working. And as AI becomes increasingly complex and opaque, it's more and more important for us to be able to understand why is it doing [03:22:01] what it's doing and how do we interpret it. And um it's useful to you know the the the idea of working on llama rather than other LLMs allows you to work with with the um sparse autoenccoders and the [03:22:14] probes in ways that you can't with other AI and and having access to the understanding of how these systems are working is going to become increasingly important. So, not just building better simulators, but understanding why [03:22:29] they're simulating in the ways that they are. And and that's going to be an issue. Um, in terms of Gabriel and the the the the paper um on on using this as a a systematic LLM tool to measure [03:22:44] qualitative data. Again, I think this points to both the enormous potential of us being able to replace a lot of wrote [03:22:56] um repeatable tasks of, you know, quantifying data, collecting it, analyzing it at at scale in ways that are going to be consistent. You know, if you employ four different RAS to go and do this and they each have a slightly [03:23:10] different encoding techniques and so forth, you might end up with data that's much worse than using a very simple um system that you can give explicit directions and you know it's following very specific rules in how it's coding [03:23:24] things. So it can do it with accuracy, it can do it directly as was mentioned and the cost replicability and so forth the scale is is unmatched. Uh I think one thing that that comes up in sort of [03:23:37] thinking about this is this approach could be used to start dealing with a problem that we've had for years which is understanding measurement error out there in data data sets that are existing and have been used quite a bit. [03:23:51] Um you know this this morning's talk about building everything on one pillar. [03:23:55] um understanding to what extent are we sort of using data which was collected by humans or constructed by humans in ways that might be somewhat idiosyncratic. So are there ways that [03:24:08] that we can use this kind of tool to better understand and collect parallel data sets um other kinds of data sets and begin to to understand measurement error better. So I think there's a lot of promise in this that goes beyond just [03:24:22] using it as a single tool for single studies but but using it in more general ways to understand measurement. [03:24:30] Okay, I want to say a few things that that I think are are sort of interesting big picture um issues that we face as AI enters into our world of of data and and so forth. And one is a lesson from a [03:24:44] data boom that came in the 1980s in finance and people you know there was suddenly minute-by-minut stock data that was available. People started analyzing it like crazy. We found the day of the week effects, the January effect, the [03:24:56] index effect, um large cap effect. There were all kinds of of particular anomalies that were coming up in sort of shiny new facts. And as we start deploying this um these tools, it's [03:25:09] possible that we can become uh you know just inundated with different interesting relationships that we couldn't have seen before. And I think it's that means that we have to start disciplining ourselves or we're just [03:25:22] going to be under a sea of of potential research questions. And so somehow managing to to understand what we should be asking, how we should we be asking it rather than just, you know, throwing it out at everything we can see is is going [03:25:36] to be hard and theory in the in in this age of AI can still be very useful in sort of guiding the deployment. What are the questions we should be asking? How do we assess um AI's use and and output? [03:25:49] Um and and how do we enrich the models in ways that are useful? [03:25:54] And I think there's um one kind of interesting paradox I've been thinking about quite a bit and and I think that this is going to become increasingly um important and and you might think of this as like the chess go paradox. So [03:26:08] it's been just about 30 years now since uh Deep Blue beat the a grandmaster. Um it's been about 10 years since uh a Go Alpha Go um beat one of the best Go players in the world. So, you know, [03:26:22] we're getting to a point where AI can do things in some ways better than humans, at least in very specific kinds of tasks. Then the question becomes at some point if it's doing things that we can't understand. So, it it generates kinds of [03:26:37] plays that look very peculiar and so forth. So, that teaches us things, but at the same time, it might be that we can, you know, it's it's doing things that are beyond our abilities to to [03:26:48] comprehend. Then how do we make use of that? Do we just trust it? So if this is, you know, if we get a simulator that does better at weather prediction, that's fine. Maybe we don't need to know exactly how why it's predicting exactly what time it's going to rain tomorrow [03:27:03] and how much it's going to rain in different parts of California. Um, that would be fine. Um, what if we start using it for predicting what's going to happen to three different potential rate decisions by the Fed and we say, "Okay, [03:27:17] here's it can give you really accurate predictions. We can't tell you why it's predicting that, but we can tell you it's doing that. So, these systems are going to be do able to do that. Um, do we want to understand why a certain rate would would lead to one thing in the [03:27:31] economy versus another? Um, and you know, eventually when trade policy, you know, understanding things like trade policy and what the implications of trade policy are, we can be using these tools to create new simulations, to [03:27:44] analyze larger and larger data sets, to build these systems. And these systems are going to become increasingly opaque. [03:27:50] And so the kinds of tools that we're talked about today to measure, understand, and again this this word mechanic or you know term mechanical interpretability, we need to be able to look under the hood eventually if we [03:28:04] want to understand the and explain why things are working. And so that I think this an interesting paradox that we're going to have that it's going to be doing some things we'll never be able to interpret and other things that we will and you know how do we handle that ## General Q&A (03:28:17 – 03:41:40) *Shared across the three papers in this session.* [03:28:17] trade-off. So thanks [applause] >> really interesting set of questions from from Matt and from the presenters. So why don't we line up over here for questions and uh while you're all doing [03:28:31] that I'll take the uh moderator's privilege to to ask the first first question but go ahead and line up there. [03:28:37] Um, so it's interesting all these papers thinking about where these things are seem incredibly powerful in many ways. [03:28:44] Where does the power come from? I was disappointed to learn that it's not magic after all, but it seems to come from the the knowledge that's in it. And I think there's some a bad aspect to that and a good aspect. One that I'd love to have people discuss a little bit more is the risk of contamination. I was [03:28:58] actually surprised at how bad the uh U u LLM did just out of the box on 1120 because the answers were right in the corpus and they could have just looked it up but but you know and sometimes they do that and you don't realize it. [03:29:11] Um, but also there's there's some potential for it to be really powerful. [03:29:15] And that also has an interesting thing is as as the technologies get more and more powerful, they're going to read all of your papers as well and learn the four steps to making a good uh simulator and know the relevant theories that it [03:29:29] should be drawing on, including some that we didn't think of in all likelihood. And then you sort of start wondering at at some point will the will the um out of the box version actually be so much better because it's it's drawing on everything and then more that that any of us can think about. So I'd [03:29:43] be curious how how you react to to those pros and cons. [03:29:48] And it's that's a question for any one of the the papers who presenters who like to Yeah. Go ahead. [03:29:56] >> Is this working? Um I think it's a great question. I'd first of all say like I'd love for my paper to be obsolete. I think that would be kind of cool. Uh I don't have a perfect answer, but I think there's a thing I've been thinking about like good [03:30:09] memorization and bad memorization for these foundation models for making predictions. And I think since it's an economist crowd, I think that good memor like bad memorization would be a student who is learning game theory for the [03:30:23] first time and memorizes the fact that defect effect is the pure strategy Nash that you should always play. like good memorization would be a student memorizing that a Nash equilibrium is when everyone is best responding to one [03:30:37] another. And I don't have a perfect answer of how I would differentiate between how to learn those two different types of memorization. I think they are in effect both memorization. Maybe like a cognitive scientist will argue with me that like one is a learning [03:30:51] memorization. I don't know how how to articulate those differences. But I think us being able to specify when models, us being able to understand those two different types of memorization better is going to be is going to be key. And I I've been thinking about it a lot. I don't have a [03:31:05] perfect answer, but I think that's the distinction. [snorts] I think that there is a difference between memorization and contamination [03:31:19] in the sense that just because a model knows something doesn't mean it actually is using it in making a decision or something like that. Um so like just because it knows of a game or knows a story um or even knows the outcome of something like we do some tests in our appendix of our paper on this like [03:31:33] doesn't mean it uses that when we're asking it do a measurement task. So I think you have to be like more careful if that like you know in a prediction task for example where that's um more likely to happen um on prior data. But um but I think that like part of what we [03:31:47] see in our paper is that like even though it could have known some of these labels or known the game or known stuff like that that um the models are you know kind of like we are capable of compartmentalizing to a certain extent. [03:31:57] We don't understand that full extent but but I think that's there. Um just to quickly add to that um in in some past work we've uh we've like actually looked for memorization when we know that the model was actually trained on the data and even then those rates of [03:32:12] memorization are maybe in the 10 to 20%. And part so I was I have I think one should be worried about memorization in a statistical way. It's kind of a case by case uh case by case question. But I think the the rubric that like just [03:32:27] because a model is trained on something is necessarily spitting things out verbatim from the text doesn't really hold partially because these models are trying to compress all the information that they're trained on. And so even if they're trained on a bunch of things, I [03:32:40] don't remember like chapter six of a book I read even though I read it in if I read that in third grade. So the models have kind of that same feature. [03:32:48] And so we have to start to get at a prompt by prompt or a case by case basis of what are they relying on when responding to a particular answer and that gets a bit bit more trickier. [03:32:58] >> Questions. >> Okay. So I have two questions for Benjamin. Um so first I wanted to know if you've tried any of the new models with your methodology like if you use GPT.5.4 four or 5.5 and if the baseline [03:33:12] improves when you use a better model and if the difference between your method and the baseline changes if the model improves and then the second question I was wondering on so you talked about the first moment on the distributions and that it it gets better I was wondering [03:33:26] on the second moment because we were talking before about kind of the the variance on the responses and how there's less variance for the interviews for example and I was wondering if the if the responses be if there's as much variance as uh there is for human [03:33:41] responses on on the AI responses or or does the variant shrink or uh if there's any comment on that. [03:33:49] >> So for the first question, I haven't tried more advanced models on our games specifically, but I have tried more advanced models doing the approach in other settings and I know other people have too. And I've I don't see a clear pattern that [03:34:03] the LLMs do a better job of predicting human responses unless human responses are like more rational. In that case, better models do tend to predict it. So, but in games like the 1120 game where it's hard to say what the most rational [03:34:18] thing is because the rational thing to do is to know how your opponent what your opponent is. Like I am rational to pick 18 if I know my opponent is a 19 picker. Um, so to that point, not not necessarily seeing it better across [03:34:32] settings and have tried the approach, although not for the specific games on other models. And then for the variance, uh, I haven't explicitly checked on the variance. I think you could kind of extend this to trying to match as many moments as possible as part of the procedure. I think that'd be like a cool [03:34:46] extension. And I think kind of the kind of what I would like to take what I would hope I would be conveying from the method is not that like one had to follow a particular optimization approach when trying to set up these uh simulations. I'm a little agnostic to [03:35:00] the optimization approach. I think there's lots of things you could try. [03:35:03] I'm guessing the more data you have, the more moments you could match probably means better, but it's an empirical question we'll have to explore. [03:35:11] >> Thanks. Okay, this is a question um about the application of a Gabriel. It's really interesting application we have here. My question is uh I want to take the author [03:35:23] stance on to what extent the post application human involvement are needed after you assemble the data. To take an example you know when we write papers there is a first stage um you know we circulate within the department in brown [03:35:38] bags and we keep getting more and more comments and all of these are you know thoughts and uh improvements is not really they go a little bit deeper and benign you know they tackle benign issues. Um so taking the example of the [03:35:53] adoption rates one may say well you know when historians look at the wave historical events we select those diffuse slowly and become important and when we think about the the what we [03:36:06] Wikipedia categorize as more uh impactful events in more recently maybe it's the faster one that got picked there. So things like that just as example uh to what extent do you guys think we can close our eyes and rest [03:36:20] easily and to what extent we still need lots of brown bags the seminars like this. Thank you. [03:36:25] >> Um well I'm happy we're doing this seminar. So uh let's keep doing that. [03:36:29] But I I think that um the uh uh I think you're it's a great question because I think that as we try to treat it in the paper is thinking of this as a measurement tool like a lot of other types of measurements like a survey or a ruler or whatever else you might use in [03:36:43] a scientific procedure which still means you need to have scientific discipline around it. Um and obviously especially because it's a newer method you want to have a particular rigor around that. So I think that like the like for example with the tech adoption stuff there's all sorts of still decisions that we have to make. maybe GPT will be able to make [03:36:58] them later on um as a research designer but we have to make about oh when do we want to cut off the data because we're afraid of skewing of that sort of adoption rate thing um what do we consider a historical significant technology what other data like um other data sets do we want to validate against [03:37:13] so all those decisions I think still need to make be made before and after I do think one shift um alluding to the um to the discussant um point is like I think that um that there is a uh a need [03:37:27] for understanding what we do with an explosion of data. Um I think that you know it just is very easy to measure a ton of attributes. So having some more um rigor and thoughtfulness about um what we measure or do we measure on like [03:37:40] a one sample and then only run it once we've done all our experimental methods on a leftout sample or things like that to to avoid like p hacking or other sorts of concerns that can that can emerge from just having so much data. [03:37:57] All right. Uh, I appreciate all these presentations. I have like a thousand questions, but I will limit myself to one. Uh, and this is for the authors of the Gabriel paper. So, um, I think the scope of the paper is very impressive, but I was kind of surprised about the [03:38:10] complete absence of references to the existing tools in NLP. In fact, like natural language processing doesn't appear in the paper at all. Um, and you know, like since 2018, we've had tools that do this with just encoder only [03:38:24] models. Um, and I was wondering if you've compared your data and accuracy to any BERT models either offtheshelf or fine-tune because you know it seems to me the gold standard in NLP is using uh [03:38:38] birectional encoders and just like that's my question and let me just say as a comment like you know these these these BERT models just run on our laptops. They don't require a server farm to to run and they basically do the same thing. And so seems to me like [03:38:52] maybe the future isn't just using these massive behemoths that we have to prompt to do the work, but to like build better small tools, small birds that that do the job and that we can just run on our computers. [03:39:07] >> Um, so um I think like in terms of comparisons like yeah, there's there's a few comparisons we do in the paper, but I think there's like more to be done. I do think part of the idea we're trying to get at here is something pretty orthogonal to what you can do with something like BERT, which is BERT or [03:39:22] some of these more recent um usages in um NLP are stuff you can train that you have to like have technical skill to train to match human labels or match human performance on one specific labeling task or a narrow range of labeling tasks. What we're trying to [03:39:36] accomplish with GBT um law LLM scale models is the general nature of human comprehension, the general idea of human labeling that you just ask the question, you don't train anything, you don't do any steps like that. And trying to [03:39:50] establish that as a competent method that that isn't to exclude that like a set of other methods might not be useful. Like for example, there's still cases where you want something really lightweight to analyze like 10 um 100 million documents and you need [03:40:03] something more lightweight than an LLM. But what we're trying to show is that LLMs are broadly capable at this sort of comprehension types task. Um and they do this labeling task um about as well as we would get from human labelers. Um [03:40:17] which is also the gold standard for BERT and other sorts of methods. So we try to show that generalizability, the technical ease for more researchers and accessibility. And then finally, um that it works for pretty small models that are still LLM grade. So a small LLMs can [03:40:32] still do this at scale. >> Okay, we're about out of time, but one quick question and quick answer. [03:40:37] >> Yeah. So for for Ben, so you showed us the statistic that uh your your more fancy model is doing better on average than the the ontheshelf model, but maybe I misunderstood it, but I didn't see [03:40:50] like how good is the model doing like in an absolute sense. [03:40:56] >> So it's it's actually a little hard to quantify. So the games are incredibly variable. Some of them are like the 511 game and some of them are the 2045 game. [03:41:05] And so like kind of doing traditional distance metrics are not like really interpretable over the vast majority of the games. Um the we have like a bunch of statistics about the optimized agents, you know, putting the majority of the probability on the choice the [03:41:20] people picked most like 80% of the time. And just in general, they're putting substantially more probability on the choices the human took most in the baseline. But honestly, it's a little bit of an open question about how to like give a really interpretable metric of how to compare so many games at once [03:41:34] with vastly different action spaces. We're still working on it. I'd love to talk about it. [03:41:39] >> Okay. Thanks. So, uh before we go to lunch, we're going to do a photo out front so future AIs will know we were all here at this moment. Um so, if you could please uh Elsa, is it right out there in in front there?