Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. =Transcript of the talk, from the video's captions. Auto-generated: speaker names in particular are unreliable. = # Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews Authors: Brian Jabarian, Luca Henkel Discussant: Kristina McElheran Video: https://www.youtube.com/watch?v=VdvT0JzwMHU&t=3807s ## Talk (01:03:27 – 01:24:09) [01:03:33] paper I'm going to present you was my drum market paper. Um it's quoted with Luca ankle from Arasmus University. Uh I think um I would like to start by basically the two first paper that we saw now were interesting because they [01:03:47] show you a bit one side of the frontier of AI which is basically AI as a way to provide information um to humans. So it can be coding, can be writing, can be any type of task. But there is another frontier that it is less um exposed for [01:04:02] now in economics but is going to become more present which is basically also seeing AI as a technology that can collect information from direct interaction with the world. So you can think of robots, you can think of like in particular Whimo, right? When you are [01:04:16] entering a car, basically Whimo, what is doing is collecting information live and adapting adapting their trajectory based on what it collects. Another side of this physical world that actually is really important in the economy is conversation like talking with humans [01:04:30] and collecting information from talking to humans directly basically. And so this technology is not just an LLM put in voice, right? It's called voice AI because it's a layer market that combine three different types of technology. one is like from speech to text recognition [01:04:45] then you put an LLM on it and then you go back to the LLM towards speech basically and you have to do this at a macro second level to have this naturalness of a conversation and it's not just um naturalness of conversation but you need also the capacity of this [01:04:59] technology to be able to adapt adapt the follow-up and the answers basically and so what we are studying here I think there is like a automatic uh if you could stop the that would be great otherwise I have to rush [01:05:13] speak faster and faster. >> Yeah. Like I can speak fast but with my with my French accent and the Yeah. [01:05:20] >> Um Yeah. Great. [laughter] All right. You're Jesus. [01:05:35] Oh my goodness. Well, it's nice. uh we don't like surprise in economics. So I mean it's a nice preview of what I'm going to present you. Um so what we are studying basically in this paper is a particular context of conversation which which [01:05:49] matter a lot in economics which is uh job interviews. Job interviews are like one of the most human intensive stage of a hiring process because you have to talk to a human being with a pro with like a goal which is collecting information from this human being. Um [01:06:03] the context in which we are studying job interviews here is um BO market. So the business process outsourcing market. So Google, Netflix and all these tech company most of the task that like involve admin legal or consumer services [01:06:18] basically delegate this to like external providers. These external providers is in in my setting my partner was teleerformance. Basically they have a recruitment process out sourcing a firm that is which is making this hiring on [01:06:30] their behalf and in this setting there is no need of like screening CV or you know resume or whatever because what you want to see here because you're recruiting for customer service position you want to see how people behave in a [01:06:44] conversation. So that's why the job interview here is like really key to this process. Um the other thing I want to say thanks for the slides. Um is that basically the BO market is a high volume market. This is another part of the [01:06:58] conversation that I don't thanks a lot for the time that I don't see um being discussed a lot in economics but actually there is a lot of gain and economic growth and we welfare improvement possible there which is basically um high volume market means [01:07:12] that the firm I work with process five million candles per year. So you can imagine the number of interviews that the human being has to conduct per per day and per month etc. And so the this volume or this scale creates a natural variance in the way the human recruiters [01:07:26] perform this job interview. And in this setting when you have high volume mark high volumes setting variance is basically a negative um feature both for the for [snorts] the employers and for the candidate. basically create like a [01:07:39] distortion the match quality simply because a humans you know sometimes they forget to ask some questions or like don't follow the expected gain etc. And so the question we're asking in this paper is what happen when you replace [01:07:51] human recruiters in the job of conducting job interviews. What happen in terms of like hiring outcome, job offers etc. But also downstream data like match quality, retention, productivity once the employees on the candidates are like hired and working as [01:08:06] an employee. How do we pursue this question? Well, we basically uh run a large scale fit experiment with 70,000 candidates with a real firm basically called page global solution the subsidiary of tele performance. In this [01:08:20] setting the candidate were randomized in mutual branch. The two ones I'm going to focus here is getting a job interview with a human recruiter or getting a job interview with an AI voice AI voice agent. So it's over the phone. And one [01:08:34] thing I really want to stress is that in this setting the the task that we are automatizing is the fact of conducting the job interview. After the job interview is conducted, this job interview is being reviewed by a by a human, a recruiter who is also taking [01:08:48] the decision. So they have access to the the audio, the voice data, but also the transcript. Basically, there is no deception involved in any way by compliance. So the candidate here they're aware of being interviewed by an AI and that someone is taking the [01:09:01] decision on the on the basis of this geometry. [01:09:06] Um okay so I will skip the literature but because there is not a lot of time but this paper basically contribute to four different strand of the literature. [01:09:14] One is that we've seen in the actually also the previous papers and like in by other people in the room as well that like augmentation like when people use AI for coding or writing ML whatever this can improve productivity but also here what we are showing is that AI [01:09:29] automation in the act of like collect information from the physical world can also improve output in our case match quality um there has been a lot of work on algorithmic screening in the HR and labor economics but this happened at two different stages of this recruitment [01:09:44] process before the job interview for like screening CV extracting information from a CV or after the job interview is done by basically making a recommendation. [01:09:53] Here we are making this uh screening in the job interview itself. Um the third thing that we are also interestingly observing in this paper is that when you automatize one part of the of the of your pipeline and then there is another [01:10:08] human being after the the humans basically how they look at the the AI signal not just as an AI recommendation but just the mere fact that the data was collected by an AI how does that change basically they're waiting on different [01:10:22] source of information to make the decision so it's not like people being averse necessarily around AI recommendation, but it's truly more about uh how to how to teach or like train humans to basically read and trust more AI signal in general. The last one [01:10:36] and it was like one of my the motivation of my work there was like I was trying to find a setting the hardest one I could find and could access to where I would show the AI would fail basically and here you have an expert task which is not just recommending or like [01:10:51] predicting but actually talking to human being just it is not like a good form of like the voice and then you have like you know a checklist you have a conversation here it's a very hard and complex task so I was thinking that basically AI would fail and it didn't so [01:11:04] that make me thinking about like what do we mean about human expertise and the boundaries of like human AI expertise I would be more cautious on you know thinking that AI will never do X um this type of claim okay so I will go briefly on the experiment the weight here of the [01:11:18] three branches that we have were negotiated with a firm um the third branch I'm not going to present in this paper but like we have a follow-up paper if you want to have a look on how choice when people are the choice between AI and recruiter can be interesting as a [01:11:32] way to uh screen people on AI immediately label But basically here the two sources of information that the human recruiter is using to make a decision are the following. One when they get the job interview either by an AI or if they [01:11:47] conduct themselves. So they have to basically give a score one to three and give an open comment uh open open-ended evaluation of the job interview. But after the job interview is done the candidate have to take an independent [01:12:00] test to measure two different skills. One is language. So like I did as a French. So you have this B2, B1, C2, etc. because you're working for US companies and the other test that they have to do is analytical score which is like how people are capable of solving problems. Remember you're a customer [01:12:15] service agent. So when people call you you're here to solve a problem and so what happened there is like a human recruiter uses two source of information and take a decision here. If the candidate is good enough according to the standard of the firm basically they issue an offer. [01:12:29] Okay. um for the last 15 minuteish I'm going to show you basically the results and then uh the mechanism for why we are seeing what we think. So the first type of result I'm showing here are like mechanical or volume like when you [clears throat] replace an AI recruiter when you replace a human recruiter by an [01:12:44] AI recruiter in the job interview what happen is that you observe a large increase in job offer rate by 12% more job offers mechanically this also goes to like more job starters but also to more retention after 30 days. So the retention here is a key metric in this [01:12:59] market because you have a huge turnover in this BPO market. So this is a main metric for match quality. I can discuss this more in in in in offline. Okay. [01:13:08] Here we see unconditional effect that also like our persistent through time because we have access to retention up to 120 days. But so far there we don't know exactly why we are seeing those retention gains, right? Is it just mechanical or is there something [01:13:22] happening in term of like how the AI is conducting a job interview? So to get there what we are looking first is the conditional effect of ma of of of of accepting a job offer and retention based on the fact of having gotten a job [01:13:36] interview AI. And what we see here is that we still have gains uh for people who got a job between AI. So 6% more job offer job starters but also conditional retention is increasing for six 30 and 60 days after it's not uh significant [01:13:50] because the sample size is shrinking um because of this turnover. So this is an indication that there might be something going going on there but we're still you know like maybe people are staying for the wrong reason on the job or you know so we want to nail this down as well and [01:14:04] this is why we look at like two other type of data that we're getting for employees one is the separation reasons are people getting more fired or living more voluntarily depending on how they got a job interview either AI or human and as you can see here there is no [01:14:18] difference in the reason for separation after 30 days so we have within 12% of population who live after 30 days No difference. The other type of data that's interesting is like we have this for a subset of the population. We when [01:14:30] we look at the core KPI of an employee as a customer service agents there is no difference as well between the two types of employees those who have been uh getting a job into an AI or human. So here you can also basically establish a minimal boundary on the quality of your [01:14:45] population and basically showing that this gain in match quality is not at the expense of productivity. So that's like where where basically we are going now to the mechanism to understand what's happening in the interview itself because this is exact this is the only [01:14:58] thing that we have vered in this experiment. Um okay so the mechanism in a nutshell is that is not AI being smarter or like more intelligent than the human but like in this volume market what's interesting [01:15:12] is that AI can basically collect information more consistently at bigger scale. So the mechanism we call it control variance because both the human and the AI both the human and the AI recruiter have basically to do the same [01:15:26] job. It's not like one the AI has a certain prompt and the human has another prompt. Here the human recruiter and the AI recruiter can basically have a job interview is a bundle of task which include one basically covering a certain [01:15:40] list of topics. In our case, we have 14 topics that when we cover that can goes from testing, you know, people abilities to solve like a problem. Here's a here's a here's a client that you face is angry. What do what do you do? How do you react? Negotiating the salary, [01:15:54] checking their attitionation risk, etc., etc. [01:15:58] The other type of task that basically involve when you do a job interview is like there is an expected guideline on how you cover these different topics. It is a conversation. So if I ask you a question about your education but you want to start about talking about your work experience that's fine the AI will [01:16:13] go with the flow basically but what will happen there is that are you coming back on track afterwards or not so that's another task the third task is the the language which is to pursue those topics basically both the human and the AI have recommended it sample of questions and [01:16:27] vocabulary that they can basically borrow to yes borrow to basically pursue those those topics and what we see here in each of these subtask is that AI is more consistently covering more topics. [01:16:39] So out of the 14 is covering more topics than the humans. It is also more likely to come back on track. So yes, it can basically start if you start talking about your education, it will be more likely and you want to talk about your work experience, it will be more likely to come back and not forget to ask the [01:16:52] question about education. And the third one here which is also interesting for customer service uh position is that AI is capable at the same time and having higher vocabulary richness within a conversation but also less variance across the interviews on the different [01:17:06] type of language they would use. And so that's interesting because um you can consistently expect the same level of vocabulary apply to different candidate across the day. Okay. So that's basically the key mechanism that we are [01:17:19] seeing in there and it does so without basically losing the ability to have this naturalness of the conversation. [01:17:24] It's not like a checklist, right? So that's like the important part because that's something that otherwise would be just mechanical. [01:17:31] Um to show you a bit how it looks like in term of like the covering topics as you can see the red is AI, gray um is human is humans. And so you see there there is like a clear shift of AI more likely to cover more topics than the [01:17:45] human being. Basically they have like in particular 50% chance of covering at least 10 of the 14 topics compared to a human who has 25% of chance of covering those topics. And so there is this notion of like you know uh how can we improve the human decisions by improving [01:17:59] recommendation but there is a lot of upstream work that also happened here which is well did you give me the information I needed to make the best decision basically here. So that's like an example of that. Um what we also see here is the expected guideline. You can [01:18:13] see on the gray here that like basically human beings are also like sometimes like u at the opposite of what we are expected in term of like guidelines in term of the order. So that's quite interesting to observe. [01:18:26] um AI is not only better than the average in all in each of these subtasks I'm showing not only better than the average recruiter which is what we care because the decision here is made by a human uh by an average human recruiter but also better than the 15 or 20% of [01:18:40] the top recruiters here which is also interesting. Uh the third type of things I want to say is that well there is more consistent and richer language used by AI for each of the topics that basically they're pursuing. Um you know there is this other channel that we don't chat [01:18:54] and don't talk a lot about in discrimination or welfare in a in in a labor place is like through language you know if you say to someone hey dude or like to someone else hi sir this is not the same experience but it's not like the human are doing this on on purpose or like being mean right it's just like [01:19:09] you're fresh in the morning and you're happy and you do your job properly but again 5 million candidates per year you know like it's it's a lot to manage as a human being and there it's there is also this notion of like incompleteness of contract that You cannot impose on human [01:19:23] recruiters like to make sure they follow like for each uh each interview 14 topics ask the expected guideline like you know you cannot design such a contract so they do the job um okay what [01:19:35] happened on the on on the candidates I just uh I forgot to mention how do we observe all of this is because as I mentioned at the beginning we have access to the job transcript so we look at like the behavior through looking at the transcript so we look at the human [01:19:49] the interviewer behavior in the transcript but also of the candidate behavior in this transcript. So what happened here is that uh okay five minute perfect. So what happened is that a human recruiter when looking at the job interview they they basically [01:20:03] looking for skills or like signals that they're going to use to decide whether to issue to extend you extend you a job offer or not. What are those different type of skills? Well for customer service positions basically the ci the [01:20:17] recruiter are looking for communication skills. Most of the time this is the most relevant type of skill that are looking for from the job interview in in addition to attrition risk. I give you two example of those type of uh skills that are looking for like signals. First [01:20:31] is interactivity as a candidate during a job interview. Do you do you show like a lot of interactivity during this conversation or you just like amorph and like kind of dead and what the human recruiter want is like to see candidate that have a high level of interactivity [01:20:45] because you know you're going to talk to your client on stream. So you want people who are like high energy and like having a lot of interactivity. The the thing you want to avoid as a human recruiter is that you want to avoid candidates who are using too much like back channel cues. So what's a back [01:20:58] channel Q is like this type of um expression like uhhuh mhm all during the conversation. This is something which is negative and if candidate basically use a lot of those type of skills, a lot of this backpically [01:21:11] used against them and not uh diminishing the probability of getting a job offer. [01:21:16] And so what we see here is that when we compare both the AI interviews and the human interviews is that the candidate who get the job interview and AI are more likely to be aligned with what the what the recruiter expected from a from [01:21:28] a candidate basically. So it means that basically in these two type of single that I mention candidate to get the job into an AI are more likely to express a lot of interactivity during the job interview and less likely to express those type of backend cues. And so at [01:21:43] this stage here, if I had just this data, you know, there might be the risk of having uh artificially increase your like perception of like who you are by lying or whatever. But because we have this independent data downstream um [01:21:56] which is the the the retention the separation reasons and um the the KPI we we see that basically there is no change on those type of variable but only better match stream. So that's why we say that basically AI here is more [01:22:10] capable of like extracting the relevant signal that basically matter for the position that we're looking for. [01:22:17] And so to conclude I am a bit early but like what I would like to to to to give here and takeaways is that AI interviews here improve outcome yes in term of job offers job starters and retention. But when you start when you look at the [01:22:31] conditional and the composition analysis, you see that actually this translate not in just more volume but like in better match in increase in match quality. And when you look at the transcript analysis, when you look at the recruiter side, what you see here is [01:22:44] that AI standardize coverage, the number of topics, the structure, but also the the richness of the vocabulary. And so this matters a lot at scale because this compression of the variance without losing the ability of having natural follow-up allows a candidate to more [01:22:58] precisely reveal who they are in the interview. And this basically helps downstream firms to have the better employees um in in in their setting. And so in one line there is that so if you [01:23:11] think about AI work so far you can think of two type of family AI as a provision of information to humans like the type of work which we saw uh just before my my talk like providing code providing writing etc. And this is great. This is [01:23:25] an important channel of productivity. But there is this other way which is really coming. The the minimal example of this is voice agents which is dynamic collection of information from the physical world. But you can think about what the robots and all these type of [01:23:39] new technology are going to be able to do in term of productivity as well. And there when you look at what happened is that it's not just because those robot are smart or whatever but they have this capacity of like reducing and like minimizing or replacing inconsistency [01:23:53] that as human being we are not able to manage at scale from the start basically. Um so yeah thank you very much. [01:24:01] [applause] >> Thank you Brian that was terrific. Um Christina Mckelin is gonna do the discussion. ## Discussant remarks (01:24:09 – 01:43:12) *Shared across the three papers in this session; Kristina McElheran discussed all three.* [01:24:12] Um, while we're waiting for Christina's slides to come up, I want to offer a special thank you to our discussants in order to kind of do a deepish dive on all the papers and um fit as many papers [01:24:26] as possible on the program. We decided to enlist one discussant for every three papers, which is a really tall order for the discussant. So, um, thank you Christina and our other discussants for agreeing to do this. And then we'll have [01:24:41] questions right after that. >> Yeah, we're gonna have 15 minutes for questions. [01:24:44] >> Okay, great. Well, thank you. Thank you for having me. Thank you for setting me this uh interesting and challenging task. Um I uh I uh I'm at the University [01:24:57] of Toronto, so I had to put a mandatory U into my labor market uh and skills session. So that's what this track is is ostensibly about. And you can see that we got some different takes on that. In [01:25:11] the time I have, I'm not going to be able to give lots of really detailed feedback to every paper. So, I'm going to try to step back a little bit, uh, think about some themes, think about some um, some thoughts I've had working [01:25:24] on similar uh, topics that maybe will be helpful for all the authors and then also for those of us who are consuming this work and and consuming it at scale and speed because [01:25:36] it's a lot. It's coming at us very fast along with the AI as the AI focused work and the AI enabled work. Um, I'm if I if if I'm feeling motivated, there might be some Gen X humor, also spelled with a U. [01:25:50] Um, so I am not a labor economist, but I play one in the JOE. Um, GenX humor. I worked on this great paper with some bonafide labor economists. Uh, came out [01:26:03] in 2023 in the journal of econometrics where we were seeking to understand the impact of digitalization on the workforce. So we're looking at worker level outcomes and seeking to really unpack uh heterogeneity in what was um [01:26:18] happening for these individual people and specifically what might be uh related to their age. So this is looking at um the the disproportionate impacts on older workers of big technological [01:26:30] change. But what I learned in this project sort of mainlining the [clears throat] literature and and working with the detailed employer employee data across the US was how [01:26:43] important firms are for understanding impacts on workers. And so this is not a shock to real labor economists who've been following you know the firming up in equality work in the QJ in 2019 and [01:26:56] um my co-authors in the journal of labor economics in 2016 where what really comes out of this is that so much of the driver of worker earnings of worker experience of technological change of inequality it's at the firm level or [01:27:11] even within firms at the um establishment level where the local context that shapes day-to-day activities of workers that gives them their incentives and their compliments [01:27:24] and their training has a massive impact. And what that is is jobs. And so one of the things that you have to start looking for when you're consuming um all of this literature is where the [01:27:38] word occupation, where the word tasks, and where the word jobs come in because they don't always mean the same thing. [01:27:44] But and I think maybe a couple of you need to check with your copy editors. [01:27:47] They don't like you repeating. And so the word work comes in when you really mean occupation or task or where you need to be very specific about whether this is is a job or this is a unit of work. And so some clarity sort of across [01:28:01] the board um would be useful to just to underscore what what what we're trying to talk about here. [01:28:09] I make this sound like I discovered this in 2023 or 2020 to 23 on that paper and it it it's not a new discovery. This is something that um I've been working on for a long time. Eric and has been working on it for a long time. Eric and [01:28:23] I have been working on it together for a long time. Understanding how important the complements to technological change are for determining adoption and the productivity impact. So I picked this one because it's the prettiest graph we [01:28:36] have. This is from our uh power of prediction paper that came out in 2021 where we're looking at the impact of predictive analytics on productivity at the establishment level. And what pops out of this when we're looking at the [01:28:50] productivity impacts of predictive analytics is how wildly uneven it is depending on the complements at the establishment level. So this is within firm variation. Predictive analytics with a lot of IT capital stock that's [01:29:03] the red is productive. without it kind of meh, right? If we're looking at skills, worker skills, if you have a skilled workforce combined with the technology, so you've got the right [01:29:15] education being selected in possibly probably not through AI back in the day um into your establishment, we get a productivity hit without not so much. [01:29:25] The green is my new obsession, which is workflow design and process design. And this is whether you have things organized in a continuous flow process where there's lots of sensors and there's lots of stability and there's [01:29:40] lots of management for continuity, low variance. Turns out data is very useful in those settings. So it's the match between that workflow design and the technology that drives the productivity gains. Without this, it like actually [01:29:54] crosses below, you know, zero. We start to see like maybe this is a misfit. um in in some part of the population and managerial capacity, management practices, the day-to-day management of [01:30:08] the establishment matters tremendously for predictive analytics. Now, this is not just old tech. I mean, so like predictive analytics is like the stone age now and we're looking at generative AI. [01:30:20] We're seeing very similar patterns when we're working this is with the MOPS data when we're looking at AI use. And so this is a theme that just sort of comes across over and over again. Um, you know, Eric's been working on this a lot. [01:30:31] This is Brunilson 2021. Uh, one of two, three, like how many did you have [laughter] there? [01:30:39] >> Um, but just to hit home the importance of the work context for shaping how this unfolds and why we really have to ask questions about what's going on with jobs, which brings me to ONET. Um so in [01:30:53] uh in the ONET data we have this um workhorse that's being used across um papers and across projects. Um I think most of you are familiar with it. If you [01:31:08] don't know this was was like the origin story was a department um of labor was helping people pick careers. And so it's like how you know what is it you like to do at the task level? let's match you into a job. Um, I think it's a wonderful [01:31:23] thing to do and it's a really rich data set, but we have to ask, you know, what is this purpose of this data and are is it fit for what we're trying to get it to do? And just as a little heads up for [01:31:36] for those of us interested in AI, now we're going to have AI classifying ONET stuff so we can use the AI generated ONET data to study AI. um that's going to be um an interesting [01:31:50] circular uh thing here. Um and so the level of analysis matters and I'm going to pick on Frey and Osborne instead of uh anybody in the room. Um [01:32:03] because this I think was kind of a seinal moment in how we think about using AI to think through job market impacts. This is at the occupational level of analysis and using computer [01:32:16] scientists to talk about what the exposure risk is at the occupation level. Uh they quickly determined that 47% of jobs were going away and the headlines the citations rolled and the headlines wrote themselves. [01:32:31] Right? So nearly half of jobs are vulnerable to automation. As a Gen Xer I really liked wage against the machine. [01:32:38] Um and and very few people stop to say, well this only works if 47% of occupations that are automated of all the jobs are the same everywhere across [01:32:53] all those occupations. And if we're mapping occupations onto jobs, you've got to have the same everywhere. And so a far less well-known study with only a thousand obser uh citations which you know still not trivial uh Melanie Arts [01:33:07] uh Terry Terry Gregory and Oric Zuran quasi immediately said what if we worry about within occupation heterogeneity in the task content of jobs and so they went in with the exact same machinery [01:33:20] and just said um you know what if we we keep everything the same. The only difference between the approaches is that we're going to use occupation level median task on one side and we're use the PAT data to ask workers what they [01:33:33] actually do in their jobs. And what they found was that the automation risk drops. Their baseline wasn't 47, it was 38, fell to nine, which is a distinctly less scary number. At least my students [01:33:45] find that when I tell them this. And the takeaway is that the occupation level assessment is kind of upward biased. And you know if we look into what is driving this I'm having trouble [01:33:59] reading this age. Uh overall we find that the automation potential is lower in jobs that require programming presenting training or influencing others. In contrast the risk of automation is higher in jobs with a high share of tasks that are related to [01:34:13] exchanging information selling or using fingers and hands. [01:34:18] So, first of all, I'm going to put a plug for reviving the Oxford comma so that we're not selling and using fingers and hands, but also to underscore the importance of thinking about jobs. So, this one tiny twist changes the takeaway [01:34:32] entirely. So, um I did a detailed literature review since 2017 and this is exactly what it looks like with the help of Chad GPT. um we have built this [01:34:46] edifice of thinking about exposure and worker implications that is all on this one narrow and I'm going to use Claude's favorite term loadbearing [01:34:58] column of onet data at the occupation level to tell us what's going to happen at workers and I'm a little bit worried about this disconnect now what's clear and I really enjoyed with the bick paper is that they take this seriously as well [01:35:12] and say you what happens if we're um a lot more careful about, you know, looking at workers. Um if we're going to use adoption statistics, I love the adoption statistics that are really [01:35:26] about the number of individuals who use the technology over the number of individuals instead of percentage of tasks or percentage of a list because percentage of a list is very hard for me to think about in terms of magnitudes. [01:35:40] So when I see percentage of occupations or percentage of tasks, I get a little bit kinky. Percentage of individuals, I can really dig into and really engages directly with sort of this chat log measurement debate and what use means. [01:35:53] As someone who's worked a long time with Eric often to find out what does it mean to use a technology, this work takes us very seriously and that's great. Um, I was a little concerned with the [01:36:06] truncation. So in the finest tradition of economic research, I decided to do some research and I dug into my uh occupational classification business teachers post-secary. I think that applies to many people in the room and [01:36:20] what you see when you look at the top 10 detailed work activities that research topics and area of expertise comes in at number 11 which is below 10 which is like sort of sad. Um but you know uh attending training sessions or [01:36:35] professional meetings comes in at 6 and evaluating student work comes in at 4. I don't use AI to evaluate my students. I had a colleague uh at another school who did that and got in trouble. Um so I don't use it for four. I would love to [01:36:49] use it for six. If it could attend meetings and training sessions for me that would be fantastic. I use it down here but I would be coded as a non-adopter and that troubles me a bit. [01:36:59] And so I I think coming, you know, just digging into a little bit more about taking this truncation more seriously. [01:37:06] Um I think we got a better story today about how important the individual fix effects are. It didn't come out so much in the paper. I think this is the story that the individual variation is very very high. And I think some of it's [01:37:20] coming through firms. I think some of it's coming through jobs. In the paper, there's a bit of a sense that these are not these are intrinsic human preferences or individual preferences. [01:37:30] Um, this comes through in sort of this plausibly exogenous argument about the IV, but it requires that nerdy people do not sort into nerdy firms in ways that also affect their home AI use. And this is like everybody I know, right? This is [01:37:45] highly selected. And I think thinking through what we can attribute to the individual versus their context could could enrich the story. And I got a little stuck on what overclassifying [01:37:56] um meant. And it it what these papers are all in conversation with each other, but some of it is is uh sort of complaining and I think it's it's useful to step back and say we're just getting [01:38:08] sort of different views on reality um or you know different measures of of the phenomenon. So with working with AI, I really enjoyed, you know, looking directly at usage. That's very intuitive. I often wondered if, you know, we could see directly like cloud [01:38:23] use and some of these other technologies we've been trying to track if that would have gotten us uh some purchase. Um I think this user goal versus AI action is distinctive and and conceptually novel, but I wasn't entirely sure what to do [01:38:37] with it. Um and the wage education results are really interesting. This is I think it'd be really interesting that this is coming out different from what we've seen and LLM classifiers are cool. [01:38:49] I'm I'm doing more work in this area myself. So I think you know this is in the theme of not just studying AI but using AI um which the other paper did as well. Um the level of analysis pet peeve really comes in here. So there's a lot of [01:39:03] places where the slate of hand happens and we think we're talking about workers, but we're actually talking about tasks in occupations. And that's I know we can't always be responsible for what people do with our research, but we [01:39:16] can try really hard to be super clear and direct that conversation in the right direction. Um there's some classification noise that I need some more careful treatment as you roll up into the rankings. some of these things [01:39:29] are within the the confidence interval and I think you probably want to do some bootstrapping. Um, and I'd like to know more on where and why correlating is you the correlation is weak with wages um, uh, an encounter to the broader [01:39:43] literature. So, I knew I'd be sort of sprinting to the end. Sorry, I'm I'm gonna be be super concise here. Um, so the third paper kind of like goes, you know, the opposite. So we're we're we're [01:39:58] looking at jobs. We're looking at at one job. And so this is magnificent in some respects and has different trade-offs in in other respects. And what we get out of this is a just really careful piece [01:40:10] of of you know causal data. This is pre-registered. It's big. This is really carefully done. Um we need more of this. [01:40:19] I don't know if we need to take these narrow studies and put them [clears throat] into macroeconomic models describing the whole economy, but that's a beef for a different day. Um, we need to learn more and more about all the different settings this can unfold [01:40:32] in very carefully recognizing that there's heterogeneity across those settings. Lots of meaningful outcomes. [01:40:39] It spans the full funnel. Um, the transcript analysis is is really interesting and I like trying to get into the mechanism so that we're not just sprinkling magic AI fairy dust on an existing workflow and seeing what comes out the [01:40:53] other side. That really figuring out what this does is is crucial and LLM classifiers are cool. Um, some concerns is that is again this kind of narrow setting with kind of a modular point solution. I think we need to be upfront [01:41:08] about where this could and could apply and not uh I think the recruiter evaluation results have two interpretations and one is really only uh dug into in the paper and being a little bit more transparent about that. [01:41:21] Um there is a bunch of stuff that's sort of failing in the AI arm and I'm not entirely buying that this is just sort of random or exogenous. There's yes I did read the papers. There's a funny [01:41:33] thing in figure A, appendix A6, where the AI interview has like more unavailable candidate, uh, that's statistically significant. So, don't say I didn't do my job. Um, [clears throat] [01:41:46] and and I think my last one is the more substantive comment here is that I do think this is an interesting and hard job, but I'm wondering how expert this specific interview job is, if reducing [01:41:58] variance and homogenizing it is is really what we want. And something that's not on my slide, but occurred to me while you were presenting, and kudos to all the presenters because I didn't have to like summarize anything. Um, is that [01:42:13] we're hearing about the firm's response to how this uh human interaction was maybe less ideal, but you can imagine that humans have all sorts of ticks and ways that we interact with people. And maybe we go off topic because that [01:42:26] really serves us in some other aspect of our life. I know that sometimes when I'm in a conversation with my kids and they go off topic, it's not so great if I push them back immediately to where they were. And so maybe thinking through the [01:42:39] the the performance metric is is potentially a bit narrow. Um I get asked all the time if I'm a techno optimist or a technimist. [01:42:48] I say yes because this is hard and this is complicated. But don't take my word for it because Wired this morning told me that using AI for just 10 minutes might make you lazy and dumb. So take it with a grain of salt. Thank you for the [01:43:03] great work. [applause] >> Thank you Christina. ## General Q&A (01:43:12 – 01:55:35) *Shared across the three papers in this session.* [01:43:12] >> Um yeah, so we have um one we have about a little more than 10 minutes for questions and comments. We have one microphone. So if you have a question, comment, just go kind of line up at that microphone there. [01:43:26] >> The standing one. Yeah, >> it's a standing one. Yeah. Behind there. [01:43:29] Yeah. Um so and please keep your remarks concise. But I am gonna give Eric the floor for the first uh question while others are >> right. While you guys line up just um so the thank you Christina, thank you the the talk. I have so many interesting [01:43:43] things to bring up on all the papers, but let me pick uh start with uh asking uh Brian and it's really follows up on Christina's last point. Um so this was an application where you're sort of standardizing, you know, minimizing the negative and there are lots of [01:43:57] applications, call centers, I know 79 workers where where having somebody substandard could really hurt the whole production process. But you could also imagine or there are also productions where you're actually trying to maximize variance. trying to when I'm looking for [01:44:10] PhD, you know, students, I'm looking for somebody who's like a rock star and like maybe I I want more variance or an entrepreneur, you know, the VCs on Sand Hill Road up here, you know, they would much rather have higher variance. So, how would that potentially change, you [01:44:25] know, the the benefits of having because it seemed like a lot of what the AI interviewer was doing was sort of standardizing and minimizing the downside. [01:44:36] >> I don't know if it's I have to go there. uh see if that works. I don't know. [01:44:40] >> Does it work? Yeah. Okay. >> Yeah, it works. [01:44:42] >> Yeah, thanks for the question. So, I agree with you. So, it depends on when the variance it's actually something that comes up uh often when I present this paper. Um I'm not saying variance is always bad. We're saying when when it's bad basically there is a case where [01:44:56] you are when should you deploy AI basically. So, if your goal is minimizing variance because it's bad, then you should do it. If even in job interviews if you're recruiting for a CEO or like an artist or you know etc you may want to have more variance in [01:45:10] the case of like judge for instance like a lot of the casistic like on the tail. [01:45:15] So you don't want to standardize your context because you can send people in jail for wrong reasons basically. So I yeah we're not saying variance is always bad. So I agree what you're saying. [01:45:27] >> Thank you. Next question. >> Thank you for the papers. I I have a question also for for you. Um do you think if each candidate was interviewed [01:45:39] by both AI and a human you would get something that is better than either? [01:45:45] >> Yes. So we have actually the answer in the followup paper. So it's called choice as signal and there what we do is we have a structural model on whether you should offer or not the choice and then we can have all this counterfactual [01:45:58] when you run a hybrid uh pipeline. So first of all actually this interesting because there is this story of I hear a lot of like automation therefore displacement not not so true in this case for instance because you can find a new comparative advantage of your [01:46:11] screeners and what you do there is like when you have two interviews one with AI one with human so it's like counterfactual you see actually improvement not only job offers and also minimizing separation so you may have [01:46:24] this this potential uh gain for all these frontline jobs where now you can specialize your recruitment ment not as if you're recruiting a CEO but like you know you can now differentiate a customer service agent for healthcare versus financial and there you will have [01:46:38] when you should have one or two interviews so yeah >> Harry so this might be an unfair question because it's not about the research and data directly it's about how we use them I talk to a lot of education and [01:46:52] training providers in community colleges and elsewhere trying to train people for jobs in health care or IT your business and they're terrified about all this. [01:47:03] All of them are. So given the state of the knowledge, given the state of the data, what's the best thing we can tell them? And that's either for any of the authors or Christina or anyone else who wants to answer. [01:47:20] >> Grab a microphone because we're we're live streaming. [01:47:23] >> No, I think it's super tricky. Um maybe I shouldn't be the one talking because this is not like Um, one of the things we saw when when I was doing my research, not just maybe picking on the the people uh I discussed [01:47:37] today was that there was this real um concern over uh the skill atrophy. So older workers were really at a loss when the new technology came in. They didn't uh keep up. They left or were pushed out [01:47:51] um or paid less. And so I think being very um comfortable working with the tech is going to be important. So the folks that are reacting to this by sort of pushing away and I I know we're kind [01:48:05] of AI positive in this room but um there's a lot of folks who are really negative on the technology and I think we're going to see this bifurcation where there's people who are just they're scared they don't understand it. [01:48:16] They don't like it. They don't like how it was trained. There's all sorts of concerns about the um environmental and and um energy implications of this and I think they're at real risk of getting left behind and we need to figure out a [01:48:30] way to make this not so divisive. >> Do do any of the other authors want to jump in on this? [01:48:41] >> Go ahead. >> Do you want to make a com comment to Gary? [01:48:45] >> Yeah. So quickly I think like there is a difference between the survey exposures that we see and people being afraid of AI and when you look at the data. So we have a survey in this setting as well and actually it relates also to some of the work when when basically people have [01:48:58] the actual experience of having AI the way they form their belief about their fear is completely different if they don't have the interview with an AI basically here. So you know um the notion of fear depend on where you get it from basically I would say [01:49:12] >> John Sable House. >> Great. Thanks. So this builds a little bit on Eric's question about the variance in the interviewing. And what I wonder what you showed us was how the AI compared to the average interviewer the elomeration of interviewers. I'd love to [01:49:25] know how the AI compared to the best interviewers. Right. So you could sort of re reverse engineer and think about, you know, try and pull out those interviewers who did a good job on the metrics, retention, etc. And then think about those charts you made that showed [01:49:40] us how they got there. Uh, and what was AI sort of more like the best interviewers or were the best interviewers even better than the AI? [01:49:48] Uh, tells us something about how uh, you know, this horse race uh, between the two uh, really c comes together. [01:49:56] >> Yeah, thanks for the question. Great question. So I mentioned this on the way when I presented that we have the same graph but like um when we compare to 15 top 15% candidate recruiters and AI is better than top 15 or top 20 depending [01:50:09] on the subtask of those recruiters. So you have this gain also when you compare the distribution. I don't think the race here is like to so that's one. The other thing I want to say is that despite those like great result from AI being better than the top 15 or best [01:50:23] recruiters in this setting in in most of the setting um the firm rely on the average recruiter for the decision making anyway. So if you were just using the rec the decision for the top 15 then I would be even more important what you [01:50:37] say but here a bit less because of this reason. The other thing is that um we are not in a we're not we not recruiting for stars here or like cos or PhDs. So here we we don't expect a shift in the distribution of the talent you're [01:50:50] recruiting but more like a shift of the average. So you know if you're not that better than like top 1% recruiters who cares in some sense here because of this volume but great point. Yes >> we have time for another um two questions. I think we have one at the [01:51:05] microphone. Uh if anyone else has a question, we are especially interested in questions for um the Microsoft paper and the the Bick paper but but please go ahead. [01:51:14] >> Thank you. So I'll make a statement and the question uh the statement is variation is not inherently bad but we spent last 30 years in education trying to stomp variation out of every process. [01:51:25] Um maybe more mundane recruiting when everybody's given exactly the same question that you have to ask in exactly the same way. So we've kind of been setting ourselves up I think for the AI uh replacement. But having said that uh [01:51:38] have you looked at uh what happens when you use AI for the initial cut because if I understand you for initial cut so uh I you don't immediately get to speak to the recruiter right you have to pass some so instead of ATS uh systems that [01:51:53] exist currently what happens if we actually have a conversation or selection because AI can do that at scale. uh if I understood your paper correctly that's not how you designed this but have you looked at uh anything that that does that [01:52:07] >> uh great question so no we have not because here the game that the firm which is why they want you to deploy AI is like they don't want to screen anyone before the interview they want to have every you know they want to give an interview to everyone as much as possible so there is no like screening [01:52:20] before that the only thing they look is like how fast do you answer to your CV or like to your sorry invitation for this next step in the process but That's it. So the screening here is a core process. Um the thing we are doing now [01:52:34] in the follow-up is actually um when you have AI and in the evaluation. So when you have an AI agents now looking at the transcript as well because of what Christ was mentioning we have this miscalibration with the recruiter. So [01:52:47] how can we basically help there and in this setting here you could see whether there is a difference between someone who get twice you know an AI interview and an AI evaluation or like a mixture basically. So that's uh I hope in six or [01:53:00] few months. Yeah. Thanks. >> Thank you. And we've got one more question. This is going to be our last question. [01:53:16] AI model question. [01:53:46] is that you can look at questions like are we seeing these workers that are using AI? Are these the ones that are that are using more? Are these the ones that are going to be less likely to be fire or lose their job in the next wave? [01:54:00] when you AI uses sort of when a recession comes all the just lay off the workers that are not using AI that's very correlated with you know what the mics what they need and so I think that [01:54:13] it seems like what you have is you can totally answer that >> okay [clears throat] so I would love to but I can't unfortunately because it is not the actual CPS right it's our online survey and the downside of these online surveys is to keep them affordable is [01:54:28] that you don't have a panel dimension So one thing we could look is at when they started their job, right? Um we could ask some questions to learn a little bit more about that, but we don't have a true panel dimension. We can only [01:54:41] ask retrospective questions. It's just if you from these online surveys, if you want to have a panel dimension, you have to pay them way more. [01:54:48] >> I see. But is it sorry just a followup like the CPS has the same problem and uh in that no what the one that you released at least for us publicly but you could ask retrospective questions like is this no to the same person were [01:55:02] you in a job last period? No. >> Yeah. Yeah. So so we can ask retrospective >> questions in the cross-section of people. [01:55:09] >> Yeah. Exactly. And so then you could go that Yeah. [01:55:12] >> So we could get at some of that. Yeah. >> Right. [01:55:17] >> Great. This is going to wrap up our last our first session. Um, but I want to thank all of our authors. These were terrific papers. Uh, and our fabulous discussant. [01:55:26] [applause] We now have a 15minute break. Please be back in your seats by 10:45 for the next session.