Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # What Work Does Generative AI Do? Authors: Alexander Bick, Adam Blandin, David J. Deming, Tyler R. Schumacher Discussant: Kristina McElheran Video: https://www.youtube.com/watch?v=VdvT0JzwMHU&t=2524s ## Talk (00:42:04 – 01:03:27) [00:42:08] >> [clears throat] >> So thanks for having us on the program. [00:42:12] This is a joint work with David Deming, Adam Planning and Tyler Schumacher who is also here. And just already going to get started without the slides, but uh the broad question right that we all [00:42:27] interested in I think >> was >> oh sorry uh so the broad question we're interested in is like what does AI do to productivity to employment? I'm not going to speak so much about that but rather get into uh some of the [00:42:41] measurement details and I think I'm very fortunate that Kieran presented in front of me because he's setting up the stage for me. Um all right so so how do we measure AI? We [00:42:54] are doing it with an online survey. Um and then throughout the talk I want to touch on three topics namely uh state of adoption across occupations and tasks. [00:43:05] um and then uh kind of what are some predictors of AI adoption and finally I'm going to talk about um some comparisons to the chat data from Microsoft but also from entropic and [00:43:17] cloth could I get the slides all right maybe I'll thank you okay so I'll talk about that let me go straight to measurement uh where do the [00:43:31] data come from um Adam and I we have been running this online survey which we call the real-time population survey since April 20. It's a 10-minute online survey. Sample size is about 5,000 people per wave, and currently we're [00:43:44] running it quarterly. Now, what is key to the survey is that we replicate the core module of the current population survey, which then allows us to weight our survey responses against the CPS. So that gives us a national representative [00:43:58] sample of the population age 16 to 18 to 64 by key demographics and work status. [00:44:05] Okay. Now one thing always of course to be concerned with these online surveys is selection based on unobservables. uh the people are taking these surveys different from the general population and we've written a bunch of papers where we I think at least for all the [00:44:19] topics including aspects of AI adoption that we were interested in uh these selection on observable is not a concern now the cool thing about the online service is it has extra information [00:44:32] relative to the CPS that we can add so since August 24 we've been asking about generative AI use and it's kind of by definition by workers it's a worker sample survey. Okay. So, I have not anything to say about if firms [00:44:45] completely automizing some processes that leave the workers out of the loop. [00:44:49] We can't speak to that. And since August 25, we have been asking about the top 10 owner detailed work activities and for which of those AI is used. How do we do that? um we have a way I think of [00:45:02] eliciting that works very well the detailed so SOC occupational code and then we can map those detailed work activities with the important scores okay um of course I mean there's going to be some measurement error no doubt [00:45:17] but I think uh what we show is that all task related patterns whether it's about performing tasks um related to adoption they're fairly they're quite stable across waves uh and they're consistent with a bunch of ONET benchmarks that I [00:45:30] don't have time to get into details but but I think it is working really well. [00:45:35] One more caveat I want to say is that our survey elicits uh intensive margin of AI usage and time savings on AI usage but just overall not on the occupational level. So what we're going to do is everything is going to be about the [00:45:48] extensive margin uh along the along the task dimension. Okay. and we felt it was really helpful to set up a conceptual framework to talk you through some of the material that I'm presenting today. [00:46:01] Okay, it's very simple, lots of simplifying assumptions, but it really helped us guiding our thinking and I hope you will find it useful. Um, there's going to be just a set of Occupations that have employment shares lambda O and they sum up to one. Okay. [00:46:15] Then there's going to be capital T tasks uh that can be performed in different occupations. [00:46:20] And this omega OT is the occupational chair that might be zero. None of us is performing surgery. Others are performing surgery. Right? And then within uh occupation all these task shares have up sum up to one. But we can [00:46:34] also look at occupa um tasks economywide task shares when we sum up a specific task across all occupations. Okay. And then we're going to have this task level output um that aggregates linearly into [00:46:47] aggregate output. Now what how do we think about genai in that framework? If you adopt it, there's going to be some productivity gain but there's going to be some adoption cost for the worker. [00:46:58] You know um and you know you adopt when the gain is exceeding the um adoption cost and that gives you then on the task level an adoption rate. And then if you assume that another simplifying [00:47:12] assumption that pre-Gen AI output was one in each occupation task pair then you can get the simple formula that the change in productivity by adopting AI is simply the task share weighted adoption rate time the productivity gains. So [00:47:27] what I want to talk now about is the adoption rate. Okay. So this these are adoption rates now on the three-digit CPS occupation level. And one thing that you can see is that 80% of occupations [00:47:41] um have an adoption rate of um have an adoption rate of at least 20%. And very few occupations come kind of close to 100%. What are the top adopters? It's computer programmers and public [00:47:54] relations specialists. Uh what are the bottom two adopters is maintenance and repair workers and fast food and counter workers. Now we can switch to the task level. This is on the DWA level and here [00:48:07] we see that about 40% of occupations have adoption rates of at least 20%. We don't have these extremely high adoption rates that we see on the occupation level because there are some you know [00:48:20] tasks that you would think they should have very high adoption rate but they also performed by some workers who just don't generally have these high adoption rates. Okay. So here I just you know show you the top two task and in light gray I show you the occupations the top [00:48:34] two occupations and it's prepare research posts and ports and design computer systems and like I mean you can see how that maps to computer programmers. There's lots of tasks that have zero adoption rate. So here I just picked two of them that have zero [00:48:48] adoption rate that match kind of closely with those occupations with the load adoption rates and you know presenting food or beverage information and menus to customers is surprisingly little AI involvement. Okay. Now what are [00:49:03] predictors of geni adoption through the lens of our framework? It's clear if you have a high productivity gain in a task then adoption is going to be more likely and we think these exposure measures provide a useful proxy for task or [00:49:17] occupational level productivity gains and Eric has contributed and we've seen the Lundo measure just in the previous talk and we're going to do something basically very similar in the first stage we're going to request adoption uh [00:49:30] on the illus alundu ceda score which was the one that assumes the highest AI capabilities And we do that at four levels of aggregation. So the first one I want to show you is just take occupational taskwide adoption rates and request that [00:49:43] alundus score go on it about 40 to 50% of the variation can be explained by them. So that I think is some validation some showing the usefulness of these measures but it also tell you there's lots of variation that cannot be [00:49:56] explained by that. Okay. And in order to get there, I first want to kind of go back to our framework and think about what are these uh adoption cost. There's like three different parts. It's like a task specific cost. Think about for [00:50:10] example legal requirements that prevent you from using AI. That could be very high cost. Or if it's not regulated at all, it could be very low. There could be worker specific cost. Uh on average, maybe in some sectors or occupations, people are more techsavvy than others or [00:50:24] feel com more comfortable with new technologies. And then there could be the synchronotic cost like I prefer using it for coding but I don't like it using it for slides. Okay. So in here I'm going to show you basically a very similar figure that we just saw before [00:50:38] but now we use our occupationwide adoption rate and on the y-axis is the exposure score. On the on the x-axis is the exposure score on the y- axis is the uh adoption rate. And you can see this strong positive correlation that generates that high R squar. But you [00:50:52] also see this massive variation for given exposure score and how how adoption rates differ. And you see exactly the same for tasks. And here I've switched now to the IWA level. [00:51:03] Okay. And now one thing kind of trying to summarize what what makes these tasks different where we have high adoption uh much higher adoption that kind of what the fitted line would predict or lower. [00:51:14] So let's first focus on those that are overpredicted. Okay. So think about these are dots beyond the fitted line. [00:51:21] These are healthcare clinical verification and compliance. You could easily see how uh task specific adoption costs uh like these costs like legal complaint legal considerations factor in [00:51:35] in low adoption rates there. And then we have also routine admin work which of course partly correlates with those other tasks um uh where maybe workers are less comfortable with using technologies. [00:51:47] And then in terms of what is underpredicted or vice versa what are these dots that are very far above the fitted line that are strategic planning technical development coordination problem solving. So a lot of tasks where people have a lot of discretion as [00:52:01] probably low kind of tasks barrier cost people that are very comfortable with tools have good education in that regard. So again this is no proof this is just trying to fix some intuition. [00:52:13] Okay now we go to the individual level. Now on the individual worker level or the worker task level, this exposure score has much less predictive power. [00:52:22] Okay, what has more predictive power are occupation fixed effects and task fixed effects. But I was really swamping everything is in individual fixed effects. Okay? So if you tell me you use it for one task, I'm way better at [00:52:37] predicting if you use it for another task than if you tell me what the exposure score is or what your occupation or even the task is. Okay? So we think that suggests a role for some type of experience. [00:52:50] Um so this is kind of a bit more lucer exercise but our hypothesis that is that genai experience lowers the cost of using genai in other domains. Okay. So uh the problem is of course if you want [00:53:04] to test that that you know if you think about what that experience is if you have already low adoption cost on the individual level uh that kind of tells you probably you have to accumulate more experience and that's going to also make it more likely to adopt for the next [00:53:18] task. So you have this positive correlation with the unobserved fixed cost. Okay. So we have three exercises that are consistent with the diffusion. [00:53:27] I'm not going to go into the details but two rely on instruments and one relies on retrospective information. Okay. Uh one is where we can really literally ask if I look at all the adoption in your other tasks and I want to predict [00:53:41] adoption for one task. Okay. Again this is like the clear endogenity issue but we have an instrument for that that is really super predictive even with the instrument of adoption in that focal task. um workers who are induced to use [00:53:54] Gen AI at work through employee encouragement, they are way more likely to also use it at home. Then we have this retrospective question where we ask whether people used AI in the past. You can use that as some form of experience [00:54:08] and that also predicts whether you're doing more tasks right now. Now any applied microeconomists in this room are going to nail me on all the you know identification. This is not airtight and I don't want to claim that. I think about like the consistency of these [00:54:22] three exercises just giving us an idea that it is consistent with this diffusion mechanism. Okay. Now I want to go to a comparison with the chat data and we already got a nice introduction to one of them. Um so what are these [00:54:36] chat data? It's just a bunch of chats right that you partition into tasks using some classifiers. Um we have the uh open AI measure using uh chat GPT. we have the entropic measure and then we have Kieran and Corus's user goal [00:54:50] measure because I think that aligns most closest is when we ask workers are you using AI for a task so I think that is that is has the closest map um but if you think about what is this uh what is [00:55:03] this chat chair it's nothing else that you identify a task conditional on AI usage right because you everyone is using AI in those chats and you just classify the task now we can construct [00:55:16] Exactly the same measurement in our data. Um by using base rule you can just you know it's the adoption rate times the task share uh in the aggregate economy divided by the overall adoption [00:55:29] rate. Okay. Now what would it take for these chatbased measures to recover the RPS style type of measure? Okay. First of all, each platform features the user base that is representative of the [00:55:43] economy's a AI adapters. So particular in the entropic and the um open AI measure, they exclude enterprise accounts. Okay, so that already gives you an idea that this is not representative. Then the other thing is [00:55:57] that each task instance generates the same number of chats. What do I mean by that? Take the entropic measure, the open AI measure where they select one specific chat, right? Right? They give you a little bit of context by looking at the 10 chats before. But if writing [00:56:12] involves a lot of back and forth versus searching for information is like just one or two interactions, that's going to screw the measure to writing because you're way more likely to sample this. [00:56:23] Now, of course, at the other hand, it also going to capture some intensive margin if AI is used more for writing. [00:56:29] Uh, and that is something that we cannot capture in our data. [00:56:33] Now I think what is the most trickiest thing is that each chat is assigned to an to the owner task that actually assists. So for most people of us in the room economists or economics professors we presumably use it for a lot for [00:56:46] coding but coding is just not a task in on for economists. So what does it mean for us? It would be conducting research that would be the activity where we use coding but that is not something coding [00:57:01] itself is not the task. So you see there's there's this room for m mclassification. Okay. Now empirically the chat data correlate I would say particular at the rank level highly uh with the with our measure but they are way more concentrated and I want to kind [00:57:16] of dive into that now for a little bit particular what I'm showing you here is the top 10 top gen AI task in each data set. So in open AI data it's written materials it's 15%. In anthropic, it's [00:57:30] designing computer systems at 16%. In the co-pilot user uh goal data, it's gathering information that's 23%. Now, in the RPS, it's directing organizational operation activities or [00:57:43] procedures and that's 4%. Okay? So, you see that was the much more concentrated. [00:57:48] The other aspect is all of these data sets have very different top tasks. [00:57:53] Okay? And they barely overlap. Now part of that could if for example reflect that uh in claude people are way more likely to be coders coders right at least traditionally it was very strong there so that could reflect [00:58:07] non-representativeness but I think it also is classifying something about the about the classifiers so the next bullet point I want to pay your attention to is um in the in this column the second to last [00:58:21] column we show you the RPS adoption rate for people who have that IWA. Okay, you see these are indeed tasks that have very high adoption rates. The next thing is I show you how many C workers in the [00:58:35] CPS feature that IWA and with the exception of the designing computer systems particularly adding written materials and gather information that is just not a common IWA for workers. Okay, [00:58:48] they just don't see that. Whereas it's that little bit obscure one in the RPS that gets so much traction. It's like 56% of workers actually have that IWA. [00:58:58] Okay. Now, uh kind of I said already that last bullet point, right? So the top task has a much lower share. So that's about the concentration, but it's also a task that is very common for workers. Now let me [00:59:12] try to summarize these patterns of disagreement. Okay. What is under represented in the chat data? It's higher order purpose oriented tasks. So think about conducting research. Okay. [00:59:25] Now what is over represented? It's basic generic applica actions that apply across many purposes and context. Okay. [00:59:33] Now we had a separate question where we ask about writing and editing documents and guess what 50% of people tell us they are doing this task. Okay. And among of those that are doing this task [00:59:45] 50% use AI for it. So that really maps nicely with actually what I've shown you before. The thing is there's nothing wrong with these chat classifications. I think they are they're getting at what people are doing right and it's great to [00:59:59] have these data but I think owned is just not the right framework to think about it. Okay. And I think that becomes even starker if we want to infer occupational composition of platform users through onet. One example is [01:00:14] entropic says like 37% of the tasks are computer systems computer programming tasks but that's becoming because of the classification because they classify so much as coding right um and and people [01:00:29] are using those occupational measures as like say oh this occupation is highly exposed but uh it's kind of I think mixing up different things and it's becoming even more important though in the light of the workforce and [01:00:41] transparency act of 2026 six where it's now about that these AI labs should share the occupational composition and the type of task people are using again this is all fine but if we try to think [01:00:54] of that through it I think that is where that is creating the the mismatch okay now let me conclude what are we showing we use novel task level data for us workers to talk about [01:01:07] the state of Gen AI adoption uh one thing that I want to emphasize because like there's this author and Thompson paper about expertise which makes stark predictions about wage changes employment share changes uh in the [01:01:21] occupation depending what type of task in terms of expertise level AI is is is doing and um I didn't have time to show that but we have figure seven in the paper um that that gets at that where we ask people about that [01:01:35] I think these exposure scores have been very useful they are very informative and and they do explain a lot of the vari ation and the occupation and task level but they also miss a lot of variation and we saw that in Curan's [01:01:48] data too uh at the worker and task level it's really these worker fixed effects that dominate and partly of that is driven by diffusion across domains but that also tells you if that is a story that you buy in that as one gets more [01:02:02] people into adopting it people that maybe have less of a inclination to use it you're also going to see more aggregate adoption and I think the chat data again fantastic because you can learn so much from them but there's also [01:02:16] some constraints on it because they kind of concentrate on these generic activity based tasks I just think it's nothing that is wrong with that it's just that onet is not the right framework and I don't think we can infer occupations of [01:02:29] platform users from that because of this classification so through the lens of onet we don't have a mapping from them who is adopting and what they adopting it okay I I hope that our survey based adoption rates while of course having [01:02:44] lots of shortcomings too you know I don't want to oversell but I think by on the occupation and task level they help us to close some gap and I view all these different efforts as like complimentary to learning about the [01:02:57] state of AI and how it making its way through the economy. Thank you. [01:03:03] [applause] Okay, our next paper is going to be voice AIMS natural field experiment on automatic [01:03:18] job. >> So yeah, my name is Brian Javarian. The ## Discussant remarks (01:24:09 – 01:43:12) *Shared across the three papers in this session; Kristina McElheran discussed all three.* [01:24:12] Um, while we're waiting for Christina's slides to come up, I want to offer a special thank you to our discussants in order to kind of do a deepish dive on all the papers and um fit as many papers [01:24:26] as possible on the program. We decided to enlist one discussant for every three papers, which is a really tall order for the discussant. So, um, thank you Christina and our other discussants for agreeing to do this. And then we'll have [01:24:41] questions right after that. >> Yeah, we're gonna have 15 minutes for questions. [01:24:44] >> Okay, great. Well, thank you. Thank you for having me. Thank you for setting me this uh interesting and challenging task. Um I uh I uh I'm at the University [01:24:57] of Toronto, so I had to put a mandatory U into my labor market uh and skills session. So that's what this track is is ostensibly about. And you can see that we got some different takes on that. In [01:25:11] the time I have, I'm not going to be able to give lots of really detailed feedback to every paper. So, I'm going to try to step back a little bit, uh, think about some themes, think about some um, some thoughts I've had working [01:25:24] on similar uh, topics that maybe will be helpful for all the authors and then also for those of us who are consuming this work and and consuming it at scale and speed because [01:25:36] it's a lot. It's coming at us very fast along with the AI as the AI focused work and the AI enabled work. Um, I'm if I if if I'm feeling motivated, there might be some Gen X humor, also spelled with a U. [01:25:50] Um, so I am not a labor economist, but I play one in the JOE. Um, GenX humor. I worked on this great paper with some bonafide labor economists. Uh, came out [01:26:03] in 2023 in the journal of econometrics where we were seeking to understand the impact of digitalization on the workforce. So we're looking at worker level outcomes and seeking to really unpack uh heterogeneity in what was um [01:26:18] happening for these individual people and specifically what might be uh related to their age. So this is looking at um the the disproportionate impacts on older workers of big technological [01:26:30] change. But what I learned in this project sort of mainlining the [clears throat] literature and and working with the detailed employer employee data across the US was how [01:26:43] important firms are for understanding impacts on workers. And so this is not a shock to real labor economists who've been following you know the firming up in equality work in the QJ in 2019 and [01:26:56] um my co-authors in the journal of labor economics in 2016 where what really comes out of this is that so much of the driver of worker earnings of worker experience of technological change of inequality it's at the firm level or [01:27:11] even within firms at the um establishment level where the local context that shapes day-to-day activities of workers that gives them their incentives and their compliments [01:27:24] and their training has a massive impact. And what that is is jobs. And so one of the things that you have to start looking for when you're consuming um all of this literature is where the [01:27:38] word occupation, where the word tasks, and where the word jobs come in because they don't always mean the same thing. [01:27:44] But and I think maybe a couple of you need to check with your copy editors. [01:27:47] They don't like you repeating. And so the word work comes in when you really mean occupation or task or where you need to be very specific about whether this is is a job or this is a unit of work. And so some clarity sort of across [01:28:01] the board um would be useful to just to underscore what what what we're trying to talk about here. [01:28:09] I make this sound like I discovered this in 2023 or 2020 to 23 on that paper and it it it's not a new discovery. This is something that um I've been working on for a long time. Eric and has been working on it for a long time. Eric and [01:28:23] I have been working on it together for a long time. Understanding how important the complements to technological change are for determining adoption and the productivity impact. So I picked this one because it's the prettiest graph we [01:28:36] have. This is from our uh power of prediction paper that came out in 2021 where we're looking at the impact of predictive analytics on productivity at the establishment level. And what pops out of this when we're looking at the [01:28:50] productivity impacts of predictive analytics is how wildly uneven it is depending on the complements at the establishment level. So this is within firm variation. Predictive analytics with a lot of IT capital stock that's [01:29:03] the red is productive. without it kind of meh, right? If we're looking at skills, worker skills, if you have a skilled workforce combined with the technology, so you've got the right [01:29:15] education being selected in possibly probably not through AI back in the day um into your establishment, we get a productivity hit without not so much. [01:29:25] The green is my new obsession, which is workflow design and process design. And this is whether you have things organized in a continuous flow process where there's lots of sensors and there's lots of stability and there's [01:29:40] lots of management for continuity, low variance. Turns out data is very useful in those settings. So it's the match between that workflow design and the technology that drives the productivity gains. Without this, it like actually [01:29:54] crosses below, you know, zero. We start to see like maybe this is a misfit. um in in some part of the population and managerial capacity, management practices, the day-to-day management of [01:30:08] the establishment matters tremendously for predictive analytics. Now, this is not just old tech. I mean, so like predictive analytics is like the stone age now and we're looking at generative AI. [01:30:20] We're seeing very similar patterns when we're working this is with the MOPS data when we're looking at AI use. And so this is a theme that just sort of comes across over and over again. Um, you know, Eric's been working on this a lot. [01:30:31] This is Brunilson 2021. Uh, one of two, three, like how many did you have [laughter] there? [01:30:39] >> Um, but just to hit home the importance of the work context for shaping how this unfolds and why we really have to ask questions about what's going on with jobs, which brings me to ONET. Um so in [01:30:53] uh in the ONET data we have this um workhorse that's being used across um papers and across projects. Um I think most of you are familiar with it. If you [01:31:08] don't know this was was like the origin story was a department um of labor was helping people pick careers. And so it's like how you know what is it you like to do at the task level? let's match you into a job. Um, I think it's a wonderful [01:31:23] thing to do and it's a really rich data set, but we have to ask, you know, what is this purpose of this data and are is it fit for what we're trying to get it to do? And just as a little heads up for [01:31:36] for those of us interested in AI, now we're going to have AI classifying ONET stuff so we can use the AI generated ONET data to study AI. um that's going to be um an interesting [01:31:50] circular uh thing here. Um and so the level of analysis matters and I'm going to pick on Frey and Osborne instead of uh anybody in the room. Um [01:32:03] because this I think was kind of a seinal moment in how we think about using AI to think through job market impacts. This is at the occupational level of analysis and using computer [01:32:16] scientists to talk about what the exposure risk is at the occupation level. Uh they quickly determined that 47% of jobs were going away and the headlines the citations rolled and the headlines wrote themselves. [01:32:31] Right? So nearly half of jobs are vulnerable to automation. As a Gen Xer I really liked wage against the machine. [01:32:38] Um and and very few people stop to say, well this only works if 47% of occupations that are automated of all the jobs are the same everywhere across [01:32:53] all those occupations. And if we're mapping occupations onto jobs, you've got to have the same everywhere. And so a far less well-known study with only a thousand obser uh citations which you know still not trivial uh Melanie Arts [01:33:07] uh Terry Terry Gregory and Oric Zuran quasi immediately said what if we worry about within occupation heterogeneity in the task content of jobs and so they went in with the exact same machinery [01:33:20] and just said um you know what if we we keep everything the same. The only difference between the approaches is that we're going to use occupation level median task on one side and we're use the PAT data to ask workers what they [01:33:33] actually do in their jobs. And what they found was that the automation risk drops. Their baseline wasn't 47, it was 38, fell to nine, which is a distinctly less scary number. At least my students [01:33:45] find that when I tell them this. And the takeaway is that the occupation level assessment is kind of upward biased. And you know if we look into what is driving this I'm having trouble [01:33:59] reading this age. Uh overall we find that the automation potential is lower in jobs that require programming presenting training or influencing others. In contrast the risk of automation is higher in jobs with a high share of tasks that are related to [01:34:13] exchanging information selling or using fingers and hands. [01:34:18] So, first of all, I'm going to put a plug for reviving the Oxford comma so that we're not selling and using fingers and hands, but also to underscore the importance of thinking about jobs. So, this one tiny twist changes the takeaway [01:34:32] entirely. So, um I did a detailed literature review since 2017 and this is exactly what it looks like with the help of Chad GPT. um we have built this [01:34:46] edifice of thinking about exposure and worker implications that is all on this one narrow and I'm going to use Claude's favorite term loadbearing [01:34:58] column of onet data at the occupation level to tell us what's going to happen at workers and I'm a little bit worried about this disconnect now what's clear and I really enjoyed with the bick paper is that they take this seriously as well [01:35:12] and say you what happens if we're um a lot more careful about, you know, looking at workers. Um if we're going to use adoption statistics, I love the adoption statistics that are really [01:35:26] about the number of individuals who use the technology over the number of individuals instead of percentage of tasks or percentage of a list because percentage of a list is very hard for me to think about in terms of magnitudes. [01:35:40] So when I see percentage of occupations or percentage of tasks, I get a little bit kinky. Percentage of individuals, I can really dig into and really engages directly with sort of this chat log measurement debate and what use means. [01:35:53] As someone who's worked a long time with Eric often to find out what does it mean to use a technology, this work takes us very seriously and that's great. Um, I was a little concerned with the [01:36:06] truncation. So in the finest tradition of economic research, I decided to do some research and I dug into my uh occupational classification business teachers post-secary. I think that applies to many people in the room and [01:36:20] what you see when you look at the top 10 detailed work activities that research topics and area of expertise comes in at number 11 which is below 10 which is like sort of sad. Um but you know uh attending training sessions or [01:36:35] professional meetings comes in at 6 and evaluating student work comes in at 4. I don't use AI to evaluate my students. I had a colleague uh at another school who did that and got in trouble. Um so I don't use it for four. I would love to [01:36:49] use it for six. If it could attend meetings and training sessions for me that would be fantastic. I use it down here but I would be coded as a non-adopter and that troubles me a bit. [01:36:59] And so I I think coming, you know, just digging into a little bit more about taking this truncation more seriously. [01:37:06] Um I think we got a better story today about how important the individual fix effects are. It didn't come out so much in the paper. I think this is the story that the individual variation is very very high. And I think some of it's [01:37:20] coming through firms. I think some of it's coming through jobs. In the paper, there's a bit of a sense that these are not these are intrinsic human preferences or individual preferences. [01:37:30] Um, this comes through in sort of this plausibly exogenous argument about the IV, but it requires that nerdy people do not sort into nerdy firms in ways that also affect their home AI use. And this is like everybody I know, right? This is [01:37:45] highly selected. And I think thinking through what we can attribute to the individual versus their context could could enrich the story. And I got a little stuck on what overclassifying [01:37:56] um meant. And it it what these papers are all in conversation with each other, but some of it is is uh sort of complaining and I think it's it's useful to step back and say we're just getting [01:38:08] sort of different views on reality um or you know different measures of of the phenomenon. So with working with AI, I really enjoyed, you know, looking directly at usage. That's very intuitive. I often wondered if, you know, we could see directly like cloud [01:38:23] use and some of these other technologies we've been trying to track if that would have gotten us uh some purchase. Um I think this user goal versus AI action is distinctive and and conceptually novel, but I wasn't entirely sure what to do [01:38:37] with it. Um and the wage education results are really interesting. This is I think it'd be really interesting that this is coming out different from what we've seen and LLM classifiers are cool. [01:38:49] I'm I'm doing more work in this area myself. So I think you know this is in the theme of not just studying AI but using AI um which the other paper did as well. Um the level of analysis pet peeve really comes in here. So there's a lot of [01:39:03] places where the slate of hand happens and we think we're talking about workers, but we're actually talking about tasks in occupations. And that's I know we can't always be responsible for what people do with our research, but we [01:39:16] can try really hard to be super clear and direct that conversation in the right direction. Um there's some classification noise that I need some more careful treatment as you roll up into the rankings. some of these things [01:39:29] are within the the confidence interval and I think you probably want to do some bootstrapping. Um, and I'd like to know more on where and why correlating is you the correlation is weak with wages um, uh, an encounter to the broader [01:39:43] literature. So, I knew I'd be sort of sprinting to the end. Sorry, I'm I'm gonna be be super concise here. Um, so the third paper kind of like goes, you know, the opposite. So we're we're we're [01:39:58] looking at jobs. We're looking at at one job. And so this is magnificent in some respects and has different trade-offs in in other respects. And what we get out of this is a just really careful piece [01:40:10] of of you know causal data. This is pre-registered. It's big. This is really carefully done. Um we need more of this. [01:40:19] I don't know if we need to take these narrow studies and put them [clears throat] into macroeconomic models describing the whole economy, but that's a beef for a different day. Um, we need to learn more and more about all the different settings this can unfold [01:40:32] in very carefully recognizing that there's heterogeneity across those settings. Lots of meaningful outcomes. [01:40:39] It spans the full funnel. Um, the transcript analysis is is really interesting and I like trying to get into the mechanism so that we're not just sprinkling magic AI fairy dust on an existing workflow and seeing what comes out the [01:40:53] other side. That really figuring out what this does is is crucial and LLM classifiers are cool. Um, some concerns is that is again this kind of narrow setting with kind of a modular point solution. I think we need to be upfront [01:41:08] about where this could and could apply and not uh I think the recruiter evaluation results have two interpretations and one is really only uh dug into in the paper and being a little bit more transparent about that. [01:41:21] Um there is a bunch of stuff that's sort of failing in the AI arm and I'm not entirely buying that this is just sort of random or exogenous. There's yes I did read the papers. There's a funny [01:41:33] thing in figure A, appendix A6, where the AI interview has like more unavailable candidate, uh, that's statistically significant. So, don't say I didn't do my job. Um, [clears throat] [01:41:46] and and I think my last one is the more substantive comment here is that I do think this is an interesting and hard job, but I'm wondering how expert this specific interview job is, if reducing [01:41:58] variance and homogenizing it is is really what we want. And something that's not on my slide, but occurred to me while you were presenting, and kudos to all the presenters because I didn't have to like summarize anything. Um, is that [01:42:13] we're hearing about the firm's response to how this uh human interaction was maybe less ideal, but you can imagine that humans have all sorts of ticks and ways that we interact with people. And maybe we go off topic because that [01:42:26] really serves us in some other aspect of our life. I know that sometimes when I'm in a conversation with my kids and they go off topic, it's not so great if I push them back immediately to where they were. And so maybe thinking through the [01:42:39] the the performance metric is is potentially a bit narrow. Um I get asked all the time if I'm a techno optimist or a technimist. [01:42:48] I say yes because this is hard and this is complicated. But don't take my word for it because Wired this morning told me that using AI for just 10 minutes might make you lazy and dumb. So take it with a grain of salt. Thank you for the [01:43:03] great work. [applause] >> Thank you Christina. ## General Q&A (01:43:12 – 01:55:35) *Shared across the three papers in this session.* [01:43:12] >> Um yeah, so we have um one we have about a little more than 10 minutes for questions and comments. We have one microphone. So if you have a question, comment, just go kind of line up at that microphone there. [01:43:26] >> The standing one. Yeah, >> it's a standing one. Yeah. Behind there. [01:43:29] Yeah. Um so and please keep your remarks concise. But I am gonna give Eric the floor for the first uh question while others are >> right. While you guys line up just um so the thank you Christina, thank you the the talk. I have so many interesting [01:43:43] things to bring up on all the papers, but let me pick uh start with uh asking uh Brian and it's really follows up on Christina's last point. Um so this was an application where you're sort of standardizing, you know, minimizing the negative and there are lots of [01:43:57] applications, call centers, I know 79 workers where where having somebody substandard could really hurt the whole production process. But you could also imagine or there are also productions where you're actually trying to maximize variance. trying to when I'm looking for [01:44:10] PhD, you know, students, I'm looking for somebody who's like a rock star and like maybe I I want more variance or an entrepreneur, you know, the VCs on Sand Hill Road up here, you know, they would much rather have higher variance. So, how would that potentially change, you [01:44:25] know, the the benefits of having because it seemed like a lot of what the AI interviewer was doing was sort of standardizing and minimizing the downside. [01:44:36] >> I don't know if it's I have to go there. uh see if that works. I don't know. [01:44:40] >> Does it work? Yeah. Okay. >> Yeah, it works. [01:44:42] >> Yeah, thanks for the question. So, I agree with you. So, it depends on when the variance it's actually something that comes up uh often when I present this paper. Um I'm not saying variance is always bad. We're saying when when it's bad basically there is a case where [01:44:56] you are when should you deploy AI basically. So, if your goal is minimizing variance because it's bad, then you should do it. If even in job interviews if you're recruiting for a CEO or like an artist or you know etc you may want to have more variance in [01:45:10] the case of like judge for instance like a lot of the casistic like on the tail. [01:45:15] So you don't want to standardize your context because you can send people in jail for wrong reasons basically. So I yeah we're not saying variance is always bad. So I agree what you're saying. [01:45:27] >> Thank you. Next question. >> Thank you for the papers. I I have a question also for for you. Um do you think if each candidate was interviewed [01:45:39] by both AI and a human you would get something that is better than either? [01:45:45] >> Yes. So we have actually the answer in the followup paper. So it's called choice as signal and there what we do is we have a structural model on whether you should offer or not the choice and then we can have all this counterfactual [01:45:58] when you run a hybrid uh pipeline. So first of all actually this interesting because there is this story of I hear a lot of like automation therefore displacement not not so true in this case for instance because you can find a new comparative advantage of your [01:46:11] screeners and what you do there is like when you have two interviews one with AI one with human so it's like counterfactual you see actually improvement not only job offers and also minimizing separation so you may have [01:46:24] this this potential uh gain for all these frontline jobs where now you can specialize your recruitment ment not as if you're recruiting a CEO but like you know you can now differentiate a customer service agent for healthcare versus financial and there you will have [01:46:38] when you should have one or two interviews so yeah >> Harry so this might be an unfair question because it's not about the research and data directly it's about how we use them I talk to a lot of education and [01:46:52] training providers in community colleges and elsewhere trying to train people for jobs in health care or IT your business and they're terrified about all this. [01:47:03] All of them are. So given the state of the knowledge, given the state of the data, what's the best thing we can tell them? And that's either for any of the authors or Christina or anyone else who wants to answer. [01:47:20] >> Grab a microphone because we're we're live streaming. [01:47:23] >> No, I think it's super tricky. Um maybe I shouldn't be the one talking because this is not like Um, one of the things we saw when when I was doing my research, not just maybe picking on the the people uh I discussed [01:47:37] today was that there was this real um concern over uh the skill atrophy. So older workers were really at a loss when the new technology came in. They didn't uh keep up. They left or were pushed out [01:47:51] um or paid less. And so I think being very um comfortable working with the tech is going to be important. So the folks that are reacting to this by sort of pushing away and I I know we're kind [01:48:05] of AI positive in this room but um there's a lot of folks who are really negative on the technology and I think we're going to see this bifurcation where there's people who are just they're scared they don't understand it. [01:48:16] They don't like it. They don't like how it was trained. There's all sorts of concerns about the um environmental and and um energy implications of this and I think they're at real risk of getting left behind and we need to figure out a [01:48:30] way to make this not so divisive. >> Do do any of the other authors want to jump in on this? [01:48:41] >> Go ahead. >> Do you want to make a com comment to Gary? [01:48:45] >> Yeah. So quickly I think like there is a difference between the survey exposures that we see and people being afraid of AI and when you look at the data. So we have a survey in this setting as well and actually it relates also to some of the work when when basically people have [01:48:58] the actual experience of having AI the way they form their belief about their fear is completely different if they don't have the interview with an AI basically here. So you know um the notion of fear depend on where you get it from basically I would say [01:49:12] >> John Sable House. >> Great. Thanks. So this builds a little bit on Eric's question about the variance in the interviewing. And what I wonder what you showed us was how the AI compared to the average interviewer the elomeration of interviewers. I'd love to [01:49:25] know how the AI compared to the best interviewers. Right. So you could sort of re reverse engineer and think about, you know, try and pull out those interviewers who did a good job on the metrics, retention, etc. And then think about those charts you made that showed [01:49:40] us how they got there. Uh, and what was AI sort of more like the best interviewers or were the best interviewers even better than the AI? [01:49:48] Uh, tells us something about how uh, you know, this horse race uh, between the two uh, really c comes together. [01:49:56] >> Yeah, thanks for the question. Great question. So I mentioned this on the way when I presented that we have the same graph but like um when we compare to 15 top 15% candidate recruiters and AI is better than top 15 or top 20 depending [01:50:09] on the subtask of those recruiters. So you have this gain also when you compare the distribution. I don't think the race here is like to so that's one. The other thing I want to say is that despite those like great result from AI being better than the top 15 or best [01:50:23] recruiters in this setting in in most of the setting um the firm rely on the average recruiter for the decision making anyway. So if you were just using the rec the decision for the top 15 then I would be even more important what you [01:50:37] say but here a bit less because of this reason. The other thing is that um we are not in a we're not we not recruiting for stars here or like cos or PhDs. So here we we don't expect a shift in the distribution of the talent you're [01:50:50] recruiting but more like a shift of the average. So you know if you're not that better than like top 1% recruiters who cares in some sense here because of this volume but great point. Yes >> we have time for another um two questions. I think we have one at the [01:51:05] microphone. Uh if anyone else has a question, we are especially interested in questions for um the Microsoft paper and the the Bick paper but but please go ahead. [01:51:14] >> Thank you. So I'll make a statement and the question uh the statement is variation is not inherently bad but we spent last 30 years in education trying to stomp variation out of every process. [01:51:25] Um maybe more mundane recruiting when everybody's given exactly the same question that you have to ask in exactly the same way. So we've kind of been setting ourselves up I think for the AI uh replacement. But having said that uh [01:51:38] have you looked at uh what happens when you use AI for the initial cut because if I understand you for initial cut so uh I you don't immediately get to speak to the recruiter right you have to pass some so instead of ATS uh systems that [01:51:53] exist currently what happens if we actually have a conversation or selection because AI can do that at scale. uh if I understood your paper correctly that's not how you designed this but have you looked at uh anything that that does that [01:52:07] >> uh great question so no we have not because here the game that the firm which is why they want you to deploy AI is like they don't want to screen anyone before the interview they want to have every you know they want to give an interview to everyone as much as possible so there is no like screening [01:52:20] before that the only thing they look is like how fast do you answer to your CV or like to your sorry invitation for this next step in the process but That's it. So the screening here is a core process. Um the thing we are doing now [01:52:34] in the follow-up is actually um when you have AI and in the evaluation. So when you have an AI agents now looking at the transcript as well because of what Christ was mentioning we have this miscalibration with the recruiter. So [01:52:47] how can we basically help there and in this setting here you could see whether there is a difference between someone who get twice you know an AI interview and an AI evaluation or like a mixture basically. So that's uh I hope in six or [01:53:00] few months. Yeah. Thanks. >> Thank you. And we've got one more question. This is going to be our last question. [01:53:16] AI model question. [01:53:46] is that you can look at questions like are we seeing these workers that are using AI? Are these the ones that are that are using more? Are these the ones that are going to be less likely to be fire or lose their job in the next wave? [01:54:00] when you AI uses sort of when a recession comes all the just lay off the workers that are not using AI that's very correlated with you know what the mics what they need and so I think that [01:54:13] it seems like what you have is you can totally answer that >> okay [clears throat] so I would love to but I can't unfortunately because it is not the actual CPS right it's our online survey and the downside of these online surveys is to keep them affordable is [01:54:28] that you don't have a panel dimension So one thing we could look is at when they started their job, right? Um we could ask some questions to learn a little bit more about that, but we don't have a true panel dimension. We can only [01:54:41] ask retrospective questions. It's just if you from these online surveys, if you want to have a panel dimension, you have to pay them way more. [01:54:48] >> I see. But is it sorry just a followup like the CPS has the same problem and uh in that no what the one that you released at least for us publicly but you could ask retrospective questions like is this no to the same person were [01:55:02] you in a job last period? No. >> Yeah. Yeah. So so we can ask retrospective >> questions in the cross-section of people. [01:55:09] >> Yeah. Exactly. And so then you could go that Yeah. [01:55:12] >> So we could get at some of that. Yeah. >> Right. [01:55:17] >> Great. This is going to wrap up our last our first session. Um, but I want to thank all of our authors. These were terrific papers. Uh, and our fabulous discussant. [01:55:26] [applause] We now have a 15minute break. Please be back in your seats by 10:45 for the next session.