Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # Working with AI: Measuring the Applicability of Generative AI to Occupations Authors: Kiran Tomlinson, Sonia Jaffe, Will Wang, Scott Counts, Siddharth Suri Discussant: Kristina McElheran Video: https://www.youtube.com/watch?v=VdvT0JzwMHU&t=1521s ## Talk (00:25:21 – 00:42:04) [00:25:25] presenting work with uh my great colleagues Sonia Jaffy will Wong Scott Counts and Sids Suri uh on measuring the applicability of generative AI to occupations. [00:25:35] So the big motivation uh for us at Microsoft was to figure out how we could use some of our our largecale AI usage data to better understand the economic effects of AI. Uh we thought this data source was very rich um given our our [00:25:49] reading of the literature on occupational AI exposure um starting with the work of Ray Osborne including work of some people in this room. uh but one thing that we were a bit dissatisfied about in this literature was that a lot of these were based on proxy measurements about alignments [00:26:04] between uh AI benchmarks and occupational tasks or patents or um maybe human or LLM labels of tasks and whether they could be assisted by AI. So our thinking was to use our observations [00:26:17] of what people are using AI for to get u kind of more uh real world grounded uh AI exposure metrics. [00:26:26] So uh using co-pilot logs we wanted to identify how LMS were applicable to work activities and then by extension to occupations. In order to answer this question we gathered a random sample of [00:26:38] uh Bing co- pilot logs from the US from 2024. [00:26:43] Uh and in order to align these to occupational tasks, uh we use a data source that I'm sure many of you are familiar with, the the ONET taxonomy, which uh just to give you a quick refresher, has about a thousand occupations and maps each one of these [00:26:57] to a collection of tasks, descriptions of of units of work that they do, which are then mapped to cross occupational um work activities at the lowest level detailed work activities and then these are clustered into intermediate work [00:27:11] activities. So for example, as a a computer and information research scientist, one of the work activities I do is analyzing scientific or applied data. [00:27:21] So uh our goal was to map conversations between users and copilot to IWAS and then use that conversation to evaluate the demonstrated capability of copilot to assist with that particular work [00:27:34] activity. So to give you an example of what this might look like, um this is a kind of anonymized version of a real conversation where a user pasted in a command line error message from some tool and the LM replied with [00:27:48] instructions for updating their configuration file. [00:27:52] And one of the things that we noticed when developing our classification approach was that it was very common for the task the user was trying to get assistance with to be different from the action that the LLM takes in response. [00:28:05] So in this case the user goal is to resolve some kind of error with their software and the AI in response then advises on uh a fix to this this error. [00:28:16] Another example that comes up very often is the user is doing some kind of learning or research task and the AI in response does a kind of teaching or explaining task. And then when we map these to work activities, we uh might map them to different things like [00:28:30] resolving a computer problem on the user side and advising on the use of technologies on the AI side or a research task on the user side and a teaching explaining task on the AI side. [00:28:42] And we think this reflects different ways in which LMS are applicable to work. So on the one hand if the user has a goal the LM can then assist them in completing that goal like assist them in doing research or in resolving their computer issue and then when we in the [00:28:55] onetonomy look at who does resolve computer problems it's things like programmers system administrators etc or they can stand in for the work of a third party that that user might otherwise have sought assistance from so [00:29:09] uh explaining technical details of products or services is something that IT and tech support workers So we use an LLM based classifier to identify the user goal work activities and the AI action work activities in [00:29:24] each conversation and these might um there might be multiple on each side. We then do a kind of uh scope of impact estimation from this conversation. So given the evidence in this conversation [00:29:38] uh what fraction of computer problems does this indicate that uh the LM is capable of assisting with uh we call this our our scope classifier and then we also do a completion classification uh was this a successful [00:29:52] conversation did the AI complete the user's goal so now uh I'll skip over a lot of the details of the methods for the sake of time but let me jump into our results. [00:30:04] uh we found that the most common work activity on the user side mostly involved uh learning type work activities like gather information obtain information about goods and services and a lot of this I think ties to the positioning of Bing Copilot at [00:30:17] that time as an add-on to the Bing search engine um I think this is something that has evolved quite a bit since then but we also see a lot of uh writing work activities and then on the AI side we see a lot of communication and teaching and explaining activities [00:30:32] like um providing information or assistance to the public and presenting research or technical info and it's very common to see a large gap between the work activities represented [00:30:44] on each side. So uh now that we know what what users are doing, which ones uh which of these work activities are most successful? We used a variety of signals to estimate this. One of them was thumbs feedback [00:30:58] from users. So they can leave a thumbs up or a thumbs down on a conversation. [00:31:02] And then we look at the the fraction of thumbs reactions on conversations mapping to a work activity that are positive as an indication of success. [00:31:12] And on the LM classification side, we have these two classifiers that I mentioned uh completion and scope. And what we found um that made us feel like our our completion classifier was capturing some real signal is that when we aggregate these signals to work [00:31:26] activities, we see a fairly high correlation between work activities that receive positive feedback and work activities that are labeled by the LLM as completed. [00:31:36] Uh we were concerned about thumbs feedback having a lot of selection bias since uh users might be much more likely to provide feedback on some types of conversations than others. So we lean more on our completion classifier. [00:31:49] But uh to give you a sense of what some of the thumbs feedback uh looks like at the work activity level here I'm showing the top um most positive thumbs feedback work activities and the least positive. [00:32:01] And among the most successful we see that they broadly cluster into uh research or information evaluation or interpretation uh work activities. For example, research laws, precedents or other legal data uh receives very [00:32:14] positive feedback and we also see writing work activities receiving positive thumbs feedback. [00:32:21] Uh on the other hand, at the low end, we see a lot of data analysis and uh image generation related activities. So these are things that people were were less happy with in in Bing Copilot. [00:32:33] Uh we also saw similar patterns with our completion and scope um results. [00:32:39] So given the these work activity level measurements, what can we say about uh occupational applicability? [00:32:46] Uh to measure this we uh compute what we call the AI applicability score summing over the work activities that an occupation does with weights according to uh weight uh ratings in ONET describing how important each task is to [00:33:00] that job. Uh we then use a a filter to filter out uh very infrequent work activities and weight by the completion and scope um scores for that work activity. [00:33:13] And here I'm showing you a box plot of uh AI applicability scores over um major occupational groups and highlighting uh information work occupations. And it's perhaps unsurprising that we see much higher AI applicability score for um [00:33:27] occupational groups that consist mainly of information work with computer and mathematical occupations at the very top and then a lot of physical occupations um at the bottom end. [00:33:38] We can take a deeper look at individual occupations and the work activities that contribute to them to see um where AI applicability is strongest. So these occupations generally fall under customer service, sales, writing etc. Uh [00:33:53] and then uh on the right hand column I'm showing the work activities that contribute the most to occupations in the left column. Uh so at the very top is edit written materials or documents. [00:34:05] Uh and then there's a lot of uh information providing explaining and writing work activities. [00:34:13] For example, uh if we look at explained technical details of products, this contributes to the AI applicability scores of customer service representatives, sales representatives, um and uh advertising sales agents. [00:34:26] And on the other hand, these editing tasks contribute to editors, um technical writers, authors, etc. [00:34:34] Uh, one caveat and disclaimer that is probably less important for this crowd than some others is that uh, our measures are are not um, indicative of what uh, occupations AI is likely to replace since all of our measurement is [00:34:47] from these um, chat logs. Um, there was some some controversy around our paper about uh, some of our our claims like this Washington Post article. Um, if you want to take a closer look at uh our [00:35:02] data, it's available online uh at this short link on GitHub. So, all of our occupational applicability scores as well as all of the work activity level measurements. [00:35:13] So, I want to briefly talk about how this compares to an earlier measure um or estimate of occupational AI exposure. [00:35:20] Uh in particular, the uh eldue at all E1 exposure from their GPTs or GPTs paper. [00:35:26] Uh overall we see fairly high correlation at the occupational level. [00:35:31] Uh but there are some some notable differences. Um for example AI applicability score is considerably higher for software developers um telemarketers, proofreaders uh and then some occupations that are fairly low on our score include things [00:35:45] like firefighting supervisors, facilities managers. Um so I think we by looking at real world AI usage we're uh getting a more complete picture than some of these um predictive estimation [00:35:57] based measures. So if you're interested in using an occupational exposure metric for some uh downstream work uh I'd recommend you you look at our our scores. [00:36:10] Uh let me briefly talk about some um some things that uh an econ might be interested in. Uh so we can for example look at how these scores uh correlate with uh wage or education requirements [00:36:23] of occupations and we see a fairly uh weak correlation both with average wage and with education requirements especially after waiting by employment at the occupation level. Uh and this is not something that [00:36:37] um previous work has always done. And indeed if we just look at raw correlation with no employment waiting uh we see a much stronger relationship. [00:36:46] And this is driven by some low wage low education requirement occupations that are very large and that have high AI applicability. And these are clustered in in sales and office and admin [00:36:59] support. Uh and we think some of this is kind of a a side effect of how granular ONET occupations are or SOC occupations are. Uh in some categories, occupations are subdivided into many groups. In [00:37:12] others, they aren't. And this leads to kind of very large differences in um employment across SOC occupations. [00:37:21] um which if we want to understand the kind of broad effect on on people um we think we'd advocate for doing kind of weighted analysis. So in general I say that the the takeaway here is that [00:37:34] there's very high variance in AI applicability across roles and that this variance is across the education and across the the wage spectrum and that any correlations in means are are much [00:37:46] smaller than than that high variance. We can also use some of our data to look at how different roles might use AI differently. uh you taking advantage of our separation between user goal and AI [00:38:01] action classifications. So we can compute our AI applicability scores using only work activities from the user goal side and only work activities from the AI action side. And our thinking there is that if we only look at user goals, this is just looking [00:38:16] at work activities that AI assists with. And if we only look at AI actions, this is only activities that AI is performing directly. [00:38:25] uh and our our thinking is that if AI is assisting with tasks relevant to an occupation, then uh it's more likely that people in those roles will engage collaboratively with AI in existing workflows. Whereas if AI is performing those tasks, then maybe they will [00:38:40] delegate uh some of their work to AI and shift the core focus of their their job to uh other tasks. And examples where we see high user goal and low AI action applicability are some of these very [00:38:54] physical occupational groups like uh food prep and serving and installation, maintenance and repair. And on the other hand, we see um kind of information work like business and financial operations [00:39:05] and um arts design, media occupations. Taking a closer look at what makes these occupations skewed towards one side or the other. Here on the left column, I've picked out some of the occupational groups that are most skewed towards user [00:39:20] goals or AI actions. And in the middle column, the work activities that contribute the most to that skew. So at the top there's more AI performance than assistance for these communication, training, information providing type [00:39:34] work activities. And at the bottom we see more assistance than performance for these physical work activities like adjusting equipment, connecting components, preparing foods, uh, and executing financial transactions. So [00:39:48] these are things where people are getting AI advice for, but the AI is not connecting lines to each other. [00:39:57] uh and some of the food and beverage uh type tasks are from people using uh LMS for uh cooking assistance and uh recipe generation. [00:40:08] I will say that some of this especially on for example the financial transaction side could change uh as we uh have more agents that have kind of um uh tool [00:40:19] access to interact with the real world. Uh, one last thing that that we can do with with our data is look at how these measures are changing over time. Since we had u data from a nine-month period, [00:40:32] we calculated all of our metrics for for each month separately and found that the main trend um this is a typo. This should say nine months. We found a broadening an increasing in diversity of the tasks that people were using AI for. [00:40:47] Uh this is reflected on the left. I'm plotting the entropy of the work activity distribution uh which we found to be increasing over time and then when we compute our AI applicability scores as a result we see uh an increase um so [00:41:01] we found that people were uh rather than concentrating their usage on a small set of activities um kind of using AI more and more broadly over time. [00:41:11] So overall I think um this data gives us some insights into where and how AI can change work. Um we find broad applicability across roles reflecting that many occupations have uh [00:41:24] information work components. Uh and we can use our AI action user goal split to separately identify how AI might assist or perform tasks. [00:41:34] Uh there are tons of of follow-up questions. Uh and I'm I'm especially excited to see the next presenter uh which uh will show us a kind of entirely different way of looking at at workplace AI impact. [00:41:49] Thank you very much. Uh excited to have your questions later and thanks for your attention. [00:41:55] [applause] >> That teed up the next paper well. So the next paper is what work does AI do? ## Discussant remarks (01:24:09 – 01:43:12) *Shared across the three papers in this session; Kristina McElheran discussed all three.* [01:24:12] Um, while we're waiting for Christina's slides to come up, I want to offer a special thank you to our discussants in order to kind of do a deepish dive on all the papers and um fit as many papers [01:24:26] as possible on the program. We decided to enlist one discussant for every three papers, which is a really tall order for the discussant. So, um, thank you Christina and our other discussants for agreeing to do this. And then we'll have [01:24:41] questions right after that. >> Yeah, we're gonna have 15 minutes for questions. [01:24:44] >> Okay, great. Well, thank you. Thank you for having me. Thank you for setting me this uh interesting and challenging task. Um I uh I uh I'm at the University [01:24:57] of Toronto, so I had to put a mandatory U into my labor market uh and skills session. So that's what this track is is ostensibly about. And you can see that we got some different takes on that. In [01:25:11] the time I have, I'm not going to be able to give lots of really detailed feedback to every paper. So, I'm going to try to step back a little bit, uh, think about some themes, think about some um, some thoughts I've had working [01:25:24] on similar uh, topics that maybe will be helpful for all the authors and then also for those of us who are consuming this work and and consuming it at scale and speed because [01:25:36] it's a lot. It's coming at us very fast along with the AI as the AI focused work and the AI enabled work. Um, I'm if I if if I'm feeling motivated, there might be some Gen X humor, also spelled with a U. [01:25:50] Um, so I am not a labor economist, but I play one in the JOE. Um, GenX humor. I worked on this great paper with some bonafide labor economists. Uh, came out [01:26:03] in 2023 in the journal of econometrics where we were seeking to understand the impact of digitalization on the workforce. So we're looking at worker level outcomes and seeking to really unpack uh heterogeneity in what was um [01:26:18] happening for these individual people and specifically what might be uh related to their age. So this is looking at um the the disproportionate impacts on older workers of big technological [01:26:30] change. But what I learned in this project sort of mainlining the [clears throat] literature and and working with the detailed employer employee data across the US was how [01:26:43] important firms are for understanding impacts on workers. And so this is not a shock to real labor economists who've been following you know the firming up in equality work in the QJ in 2019 and [01:26:56] um my co-authors in the journal of labor economics in 2016 where what really comes out of this is that so much of the driver of worker earnings of worker experience of technological change of inequality it's at the firm level or [01:27:11] even within firms at the um establishment level where the local context that shapes day-to-day activities of workers that gives them their incentives and their compliments [01:27:24] and their training has a massive impact. And what that is is jobs. And so one of the things that you have to start looking for when you're consuming um all of this literature is where the [01:27:38] word occupation, where the word tasks, and where the word jobs come in because they don't always mean the same thing. [01:27:44] But and I think maybe a couple of you need to check with your copy editors. [01:27:47] They don't like you repeating. And so the word work comes in when you really mean occupation or task or where you need to be very specific about whether this is is a job or this is a unit of work. And so some clarity sort of across [01:28:01] the board um would be useful to just to underscore what what what we're trying to talk about here. [01:28:09] I make this sound like I discovered this in 2023 or 2020 to 23 on that paper and it it it's not a new discovery. This is something that um I've been working on for a long time. Eric and has been working on it for a long time. Eric and [01:28:23] I have been working on it together for a long time. Understanding how important the complements to technological change are for determining adoption and the productivity impact. So I picked this one because it's the prettiest graph we [01:28:36] have. This is from our uh power of prediction paper that came out in 2021 where we're looking at the impact of predictive analytics on productivity at the establishment level. And what pops out of this when we're looking at the [01:28:50] productivity impacts of predictive analytics is how wildly uneven it is depending on the complements at the establishment level. So this is within firm variation. Predictive analytics with a lot of IT capital stock that's [01:29:03] the red is productive. without it kind of meh, right? If we're looking at skills, worker skills, if you have a skilled workforce combined with the technology, so you've got the right [01:29:15] education being selected in possibly probably not through AI back in the day um into your establishment, we get a productivity hit without not so much. [01:29:25] The green is my new obsession, which is workflow design and process design. And this is whether you have things organized in a continuous flow process where there's lots of sensors and there's lots of stability and there's [01:29:40] lots of management for continuity, low variance. Turns out data is very useful in those settings. So it's the match between that workflow design and the technology that drives the productivity gains. Without this, it like actually [01:29:54] crosses below, you know, zero. We start to see like maybe this is a misfit. um in in some part of the population and managerial capacity, management practices, the day-to-day management of [01:30:08] the establishment matters tremendously for predictive analytics. Now, this is not just old tech. I mean, so like predictive analytics is like the stone age now and we're looking at generative AI. [01:30:20] We're seeing very similar patterns when we're working this is with the MOPS data when we're looking at AI use. And so this is a theme that just sort of comes across over and over again. Um, you know, Eric's been working on this a lot. [01:30:31] This is Brunilson 2021. Uh, one of two, three, like how many did you have [laughter] there? [01:30:39] >> Um, but just to hit home the importance of the work context for shaping how this unfolds and why we really have to ask questions about what's going on with jobs, which brings me to ONET. Um so in [01:30:53] uh in the ONET data we have this um workhorse that's being used across um papers and across projects. Um I think most of you are familiar with it. If you [01:31:08] don't know this was was like the origin story was a department um of labor was helping people pick careers. And so it's like how you know what is it you like to do at the task level? let's match you into a job. Um, I think it's a wonderful [01:31:23] thing to do and it's a really rich data set, but we have to ask, you know, what is this purpose of this data and are is it fit for what we're trying to get it to do? And just as a little heads up for [01:31:36] for those of us interested in AI, now we're going to have AI classifying ONET stuff so we can use the AI generated ONET data to study AI. um that's going to be um an interesting [01:31:50] circular uh thing here. Um and so the level of analysis matters and I'm going to pick on Frey and Osborne instead of uh anybody in the room. Um [01:32:03] because this I think was kind of a seinal moment in how we think about using AI to think through job market impacts. This is at the occupational level of analysis and using computer [01:32:16] scientists to talk about what the exposure risk is at the occupation level. Uh they quickly determined that 47% of jobs were going away and the headlines the citations rolled and the headlines wrote themselves. [01:32:31] Right? So nearly half of jobs are vulnerable to automation. As a Gen Xer I really liked wage against the machine. [01:32:38] Um and and very few people stop to say, well this only works if 47% of occupations that are automated of all the jobs are the same everywhere across [01:32:53] all those occupations. And if we're mapping occupations onto jobs, you've got to have the same everywhere. And so a far less well-known study with only a thousand obser uh citations which you know still not trivial uh Melanie Arts [01:33:07] uh Terry Terry Gregory and Oric Zuran quasi immediately said what if we worry about within occupation heterogeneity in the task content of jobs and so they went in with the exact same machinery [01:33:20] and just said um you know what if we we keep everything the same. The only difference between the approaches is that we're going to use occupation level median task on one side and we're use the PAT data to ask workers what they [01:33:33] actually do in their jobs. And what they found was that the automation risk drops. Their baseline wasn't 47, it was 38, fell to nine, which is a distinctly less scary number. At least my students [01:33:45] find that when I tell them this. And the takeaway is that the occupation level assessment is kind of upward biased. And you know if we look into what is driving this I'm having trouble [01:33:59] reading this age. Uh overall we find that the automation potential is lower in jobs that require programming presenting training or influencing others. In contrast the risk of automation is higher in jobs with a high share of tasks that are related to [01:34:13] exchanging information selling or using fingers and hands. [01:34:18] So, first of all, I'm going to put a plug for reviving the Oxford comma so that we're not selling and using fingers and hands, but also to underscore the importance of thinking about jobs. So, this one tiny twist changes the takeaway [01:34:32] entirely. So, um I did a detailed literature review since 2017 and this is exactly what it looks like with the help of Chad GPT. um we have built this [01:34:46] edifice of thinking about exposure and worker implications that is all on this one narrow and I'm going to use Claude's favorite term loadbearing [01:34:58] column of onet data at the occupation level to tell us what's going to happen at workers and I'm a little bit worried about this disconnect now what's clear and I really enjoyed with the bick paper is that they take this seriously as well [01:35:12] and say you what happens if we're um a lot more careful about, you know, looking at workers. Um if we're going to use adoption statistics, I love the adoption statistics that are really [01:35:26] about the number of individuals who use the technology over the number of individuals instead of percentage of tasks or percentage of a list because percentage of a list is very hard for me to think about in terms of magnitudes. [01:35:40] So when I see percentage of occupations or percentage of tasks, I get a little bit kinky. Percentage of individuals, I can really dig into and really engages directly with sort of this chat log measurement debate and what use means. [01:35:53] As someone who's worked a long time with Eric often to find out what does it mean to use a technology, this work takes us very seriously and that's great. Um, I was a little concerned with the [01:36:06] truncation. So in the finest tradition of economic research, I decided to do some research and I dug into my uh occupational classification business teachers post-secary. I think that applies to many people in the room and [01:36:20] what you see when you look at the top 10 detailed work activities that research topics and area of expertise comes in at number 11 which is below 10 which is like sort of sad. Um but you know uh attending training sessions or [01:36:35] professional meetings comes in at 6 and evaluating student work comes in at 4. I don't use AI to evaluate my students. I had a colleague uh at another school who did that and got in trouble. Um so I don't use it for four. I would love to [01:36:49] use it for six. If it could attend meetings and training sessions for me that would be fantastic. I use it down here but I would be coded as a non-adopter and that troubles me a bit. [01:36:59] And so I I think coming, you know, just digging into a little bit more about taking this truncation more seriously. [01:37:06] Um I think we got a better story today about how important the individual fix effects are. It didn't come out so much in the paper. I think this is the story that the individual variation is very very high. And I think some of it's [01:37:20] coming through firms. I think some of it's coming through jobs. In the paper, there's a bit of a sense that these are not these are intrinsic human preferences or individual preferences. [01:37:30] Um, this comes through in sort of this plausibly exogenous argument about the IV, but it requires that nerdy people do not sort into nerdy firms in ways that also affect their home AI use. And this is like everybody I know, right? This is [01:37:45] highly selected. And I think thinking through what we can attribute to the individual versus their context could could enrich the story. And I got a little stuck on what overclassifying [01:37:56] um meant. And it it what these papers are all in conversation with each other, but some of it is is uh sort of complaining and I think it's it's useful to step back and say we're just getting [01:38:08] sort of different views on reality um or you know different measures of of the phenomenon. So with working with AI, I really enjoyed, you know, looking directly at usage. That's very intuitive. I often wondered if, you know, we could see directly like cloud [01:38:23] use and some of these other technologies we've been trying to track if that would have gotten us uh some purchase. Um I think this user goal versus AI action is distinctive and and conceptually novel, but I wasn't entirely sure what to do [01:38:37] with it. Um and the wage education results are really interesting. This is I think it'd be really interesting that this is coming out different from what we've seen and LLM classifiers are cool. [01:38:49] I'm I'm doing more work in this area myself. So I think you know this is in the theme of not just studying AI but using AI um which the other paper did as well. Um the level of analysis pet peeve really comes in here. So there's a lot of [01:39:03] places where the slate of hand happens and we think we're talking about workers, but we're actually talking about tasks in occupations. And that's I know we can't always be responsible for what people do with our research, but we [01:39:16] can try really hard to be super clear and direct that conversation in the right direction. Um there's some classification noise that I need some more careful treatment as you roll up into the rankings. some of these things [01:39:29] are within the the confidence interval and I think you probably want to do some bootstrapping. Um, and I'd like to know more on where and why correlating is you the correlation is weak with wages um, uh, an encounter to the broader [01:39:43] literature. So, I knew I'd be sort of sprinting to the end. Sorry, I'm I'm gonna be be super concise here. Um, so the third paper kind of like goes, you know, the opposite. So we're we're we're [01:39:58] looking at jobs. We're looking at at one job. And so this is magnificent in some respects and has different trade-offs in in other respects. And what we get out of this is a just really careful piece [01:40:10] of of you know causal data. This is pre-registered. It's big. This is really carefully done. Um we need more of this. [01:40:19] I don't know if we need to take these narrow studies and put them [clears throat] into macroeconomic models describing the whole economy, but that's a beef for a different day. Um, we need to learn more and more about all the different settings this can unfold [01:40:32] in very carefully recognizing that there's heterogeneity across those settings. Lots of meaningful outcomes. [01:40:39] It spans the full funnel. Um, the transcript analysis is is really interesting and I like trying to get into the mechanism so that we're not just sprinkling magic AI fairy dust on an existing workflow and seeing what comes out the [01:40:53] other side. That really figuring out what this does is is crucial and LLM classifiers are cool. Um, some concerns is that is again this kind of narrow setting with kind of a modular point solution. I think we need to be upfront [01:41:08] about where this could and could apply and not uh I think the recruiter evaluation results have two interpretations and one is really only uh dug into in the paper and being a little bit more transparent about that. [01:41:21] Um there is a bunch of stuff that's sort of failing in the AI arm and I'm not entirely buying that this is just sort of random or exogenous. There's yes I did read the papers. There's a funny [01:41:33] thing in figure A, appendix A6, where the AI interview has like more unavailable candidate, uh, that's statistically significant. So, don't say I didn't do my job. Um, [clears throat] [01:41:46] and and I think my last one is the more substantive comment here is that I do think this is an interesting and hard job, but I'm wondering how expert this specific interview job is, if reducing [01:41:58] variance and homogenizing it is is really what we want. And something that's not on my slide, but occurred to me while you were presenting, and kudos to all the presenters because I didn't have to like summarize anything. Um, is that [01:42:13] we're hearing about the firm's response to how this uh human interaction was maybe less ideal, but you can imagine that humans have all sorts of ticks and ways that we interact with people. And maybe we go off topic because that [01:42:26] really serves us in some other aspect of our life. I know that sometimes when I'm in a conversation with my kids and they go off topic, it's not so great if I push them back immediately to where they were. And so maybe thinking through the [01:42:39] the the performance metric is is potentially a bit narrow. Um I get asked all the time if I'm a techno optimist or a technimist. [01:42:48] I say yes because this is hard and this is complicated. But don't take my word for it because Wired this morning told me that using AI for just 10 minutes might make you lazy and dumb. So take it with a grain of salt. Thank you for the [01:43:03] great work. [applause] >> Thank you Christina. ## General Q&A (01:43:12 – 01:55:35) *Shared across the three papers in this session.* [01:43:12] >> Um yeah, so we have um one we have about a little more than 10 minutes for questions and comments. We have one microphone. So if you have a question, comment, just go kind of line up at that microphone there. [01:43:26] >> The standing one. Yeah, >> it's a standing one. Yeah. Behind there. [01:43:29] Yeah. Um so and please keep your remarks concise. But I am gonna give Eric the floor for the first uh question while others are >> right. While you guys line up just um so the thank you Christina, thank you the the talk. I have so many interesting [01:43:43] things to bring up on all the papers, but let me pick uh start with uh asking uh Brian and it's really follows up on Christina's last point. Um so this was an application where you're sort of standardizing, you know, minimizing the negative and there are lots of [01:43:57] applications, call centers, I know 79 workers where where having somebody substandard could really hurt the whole production process. But you could also imagine or there are also productions where you're actually trying to maximize variance. trying to when I'm looking for [01:44:10] PhD, you know, students, I'm looking for somebody who's like a rock star and like maybe I I want more variance or an entrepreneur, you know, the VCs on Sand Hill Road up here, you know, they would much rather have higher variance. So, how would that potentially change, you [01:44:25] know, the the benefits of having because it seemed like a lot of what the AI interviewer was doing was sort of standardizing and minimizing the downside. [01:44:36] >> I don't know if it's I have to go there. uh see if that works. I don't know. [01:44:40] >> Does it work? Yeah. Okay. >> Yeah, it works. [01:44:42] >> Yeah, thanks for the question. So, I agree with you. So, it depends on when the variance it's actually something that comes up uh often when I present this paper. Um I'm not saying variance is always bad. We're saying when when it's bad basically there is a case where [01:44:56] you are when should you deploy AI basically. So, if your goal is minimizing variance because it's bad, then you should do it. If even in job interviews if you're recruiting for a CEO or like an artist or you know etc you may want to have more variance in [01:45:10] the case of like judge for instance like a lot of the casistic like on the tail. [01:45:15] So you don't want to standardize your context because you can send people in jail for wrong reasons basically. So I yeah we're not saying variance is always bad. So I agree what you're saying. [01:45:27] >> Thank you. Next question. >> Thank you for the papers. I I have a question also for for you. Um do you think if each candidate was interviewed [01:45:39] by both AI and a human you would get something that is better than either? [01:45:45] >> Yes. So we have actually the answer in the followup paper. So it's called choice as signal and there what we do is we have a structural model on whether you should offer or not the choice and then we can have all this counterfactual [01:45:58] when you run a hybrid uh pipeline. So first of all actually this interesting because there is this story of I hear a lot of like automation therefore displacement not not so true in this case for instance because you can find a new comparative advantage of your [01:46:11] screeners and what you do there is like when you have two interviews one with AI one with human so it's like counterfactual you see actually improvement not only job offers and also minimizing separation so you may have [01:46:24] this this potential uh gain for all these frontline jobs where now you can specialize your recruitment ment not as if you're recruiting a CEO but like you know you can now differentiate a customer service agent for healthcare versus financial and there you will have [01:46:38] when you should have one or two interviews so yeah >> Harry so this might be an unfair question because it's not about the research and data directly it's about how we use them I talk to a lot of education and [01:46:52] training providers in community colleges and elsewhere trying to train people for jobs in health care or IT your business and they're terrified about all this. [01:47:03] All of them are. So given the state of the knowledge, given the state of the data, what's the best thing we can tell them? And that's either for any of the authors or Christina or anyone else who wants to answer. [01:47:20] >> Grab a microphone because we're we're live streaming. [01:47:23] >> No, I think it's super tricky. Um maybe I shouldn't be the one talking because this is not like Um, one of the things we saw when when I was doing my research, not just maybe picking on the the people uh I discussed [01:47:37] today was that there was this real um concern over uh the skill atrophy. So older workers were really at a loss when the new technology came in. They didn't uh keep up. They left or were pushed out [01:47:51] um or paid less. And so I think being very um comfortable working with the tech is going to be important. So the folks that are reacting to this by sort of pushing away and I I know we're kind [01:48:05] of AI positive in this room but um there's a lot of folks who are really negative on the technology and I think we're going to see this bifurcation where there's people who are just they're scared they don't understand it. [01:48:16] They don't like it. They don't like how it was trained. There's all sorts of concerns about the um environmental and and um energy implications of this and I think they're at real risk of getting left behind and we need to figure out a [01:48:30] way to make this not so divisive. >> Do do any of the other authors want to jump in on this? [01:48:41] >> Go ahead. >> Do you want to make a com comment to Gary? [01:48:45] >> Yeah. So quickly I think like there is a difference between the survey exposures that we see and people being afraid of AI and when you look at the data. So we have a survey in this setting as well and actually it relates also to some of the work when when basically people have [01:48:58] the actual experience of having AI the way they form their belief about their fear is completely different if they don't have the interview with an AI basically here. So you know um the notion of fear depend on where you get it from basically I would say [01:49:12] >> John Sable House. >> Great. Thanks. So this builds a little bit on Eric's question about the variance in the interviewing. And what I wonder what you showed us was how the AI compared to the average interviewer the elomeration of interviewers. I'd love to [01:49:25] know how the AI compared to the best interviewers. Right. So you could sort of re reverse engineer and think about, you know, try and pull out those interviewers who did a good job on the metrics, retention, etc. And then think about those charts you made that showed [01:49:40] us how they got there. Uh, and what was AI sort of more like the best interviewers or were the best interviewers even better than the AI? [01:49:48] Uh, tells us something about how uh, you know, this horse race uh, between the two uh, really c comes together. [01:49:56] >> Yeah, thanks for the question. Great question. So I mentioned this on the way when I presented that we have the same graph but like um when we compare to 15 top 15% candidate recruiters and AI is better than top 15 or top 20 depending [01:50:09] on the subtask of those recruiters. So you have this gain also when you compare the distribution. I don't think the race here is like to so that's one. The other thing I want to say is that despite those like great result from AI being better than the top 15 or best [01:50:23] recruiters in this setting in in most of the setting um the firm rely on the average recruiter for the decision making anyway. So if you were just using the rec the decision for the top 15 then I would be even more important what you [01:50:37] say but here a bit less because of this reason. The other thing is that um we are not in a we're not we not recruiting for stars here or like cos or PhDs. So here we we don't expect a shift in the distribution of the talent you're [01:50:50] recruiting but more like a shift of the average. So you know if you're not that better than like top 1% recruiters who cares in some sense here because of this volume but great point. Yes >> we have time for another um two questions. I think we have one at the [01:51:05] microphone. Uh if anyone else has a question, we are especially interested in questions for um the Microsoft paper and the the Bick paper but but please go ahead. [01:51:14] >> Thank you. So I'll make a statement and the question uh the statement is variation is not inherently bad but we spent last 30 years in education trying to stomp variation out of every process. [01:51:25] Um maybe more mundane recruiting when everybody's given exactly the same question that you have to ask in exactly the same way. So we've kind of been setting ourselves up I think for the AI uh replacement. But having said that uh [01:51:38] have you looked at uh what happens when you use AI for the initial cut because if I understand you for initial cut so uh I you don't immediately get to speak to the recruiter right you have to pass some so instead of ATS uh systems that [01:51:53] exist currently what happens if we actually have a conversation or selection because AI can do that at scale. uh if I understood your paper correctly that's not how you designed this but have you looked at uh anything that that does that [01:52:07] >> uh great question so no we have not because here the game that the firm which is why they want you to deploy AI is like they don't want to screen anyone before the interview they want to have every you know they want to give an interview to everyone as much as possible so there is no like screening [01:52:20] before that the only thing they look is like how fast do you answer to your CV or like to your sorry invitation for this next step in the process but That's it. So the screening here is a core process. Um the thing we are doing now [01:52:34] in the follow-up is actually um when you have AI and in the evaluation. So when you have an AI agents now looking at the transcript as well because of what Christ was mentioning we have this miscalibration with the recruiter. So [01:52:47] how can we basically help there and in this setting here you could see whether there is a difference between someone who get twice you know an AI interview and an AI evaluation or like a mixture basically. So that's uh I hope in six or [01:53:00] few months. Yeah. Thanks. >> Thank you. And we've got one more question. This is going to be our last question. [01:53:16] AI model question. [01:53:46] is that you can look at questions like are we seeing these workers that are using AI? Are these the ones that are that are using more? Are these the ones that are going to be less likely to be fire or lose their job in the next wave? [01:54:00] when you AI uses sort of when a recession comes all the just lay off the workers that are not using AI that's very correlated with you know what the mics what they need and so I think that [01:54:13] it seems like what you have is you can totally answer that >> okay [clears throat] so I would love to but I can't unfortunately because it is not the actual CPS right it's our online survey and the downside of these online surveys is to keep them affordable is [01:54:28] that you don't have a panel dimension So one thing we could look is at when they started their job, right? Um we could ask some questions to learn a little bit more about that, but we don't have a true panel dimension. We can only [01:54:41] ask retrospective questions. It's just if you from these online surveys, if you want to have a panel dimension, you have to pay them way more. [01:54:48] >> I see. But is it sorry just a followup like the CPS has the same problem and uh in that no what the one that you released at least for us publicly but you could ask retrospective questions like is this no to the same person were [01:55:02] you in a job last period? No. >> Yeah. Yeah. So so we can ask retrospective >> questions in the cross-section of people. [01:55:09] >> Yeah. Exactly. And so then you could go that Yeah. [01:55:12] >> So we could get at some of that. Yeah. >> Right. [01:55:17] >> Great. This is going to wrap up our last our first session. Um, but I want to thank all of our authors. These were terrific papers. Uh, and our fabulous discussant. [01:55:26] [applause] We now have a 15minute break. Please be back in your seats by 10:45 for the next session.