Notes on:

Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews

Brian Jabarian & Luca Henkel
arXiv:2607.28222
7 May 2026
generative AI · field experiment · hiring · labour markets
Talk · Paper · doi · Transcript
Made with AI: Opus 5 (reading and writing)

Part of AI and Economic Measurement, Spring 2026

“Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews,” by Brian Jabarian (Chicago Booth) and Luca Henkel (Erasmus University Rotterdam). Presented at the NBER conference on AI and Economic Measurement, Stanford, 7 May 2026; discussant Kristina McElheran (University of Toronto), who took all three papers in the session together. Written from the arXiv draft.

You run a recruitment-process outsourcing firm in the Philippines. You hire customer-service representatives for large American and European clients, at salaries between Php 16,000 and Php 25,000 a month, which is roughly $280 to $435. Your industry loses something like 60% of its workforce every year, so hiring is not a project you complete; it is a machine you run continuously. Your firm, Jabarian told the room, processes “five million candles per year” — candidates — and you can hear in the slip how large that number is even to the person saying it.

Because of the volume, you have made a deliberate and slightly startling choice: you do essentially no CV screening. The interview is the screen. Ten to twenty minutes on the phone or in a room, and the recruiter’s job is to find out whether this person can talk to an angry stranger for eight hours a day. To make that repeatable you have written guidelines — up to 14 topics, a recommended order, sample questions — and you have handed them to your recruiters with a great deal of latitude, because a job interview is a conversation and a conversation cannot be a checklist.

And here is the thing you cannot fix, which the paper is really about. You can write the guideline. You cannot write the contract. There is no enforceable term that says the recruiter will cover topic eleven at 4pm on their fortieth interview of the day when the applicant is talking about something else. Jabarian put it exactly this way on stage: “there is also this notion of like incompleteness of contract that you cannot impose on human recruiters.” So execution varies, between recruiters and within them, and that variance goes into the signal your hiring decision runs on as noise. The interesting claim of this paper is not that AI is smart. It is that an AI voice agent is a technology for making an incomplete contract complete — you can’t discipline a tired human into the manual, but you can prompt a machine into it — and that in a high-volume market this is worth real money.

The design, which is the whole ballgame. Anyway, 70,884 applications arrive between March and June 2025, of which 67,056 clear a minimal eligibility bar and get randomised. One arm is interviewed by a human recruiter. One arm is interviewed by an AI voice agent over the phone, from the same physical location, with the agent disclosing at the top of the call that it is an AI and that a human will review the interview and make the decision. (A third arm lets the applicant choose; that one is held back for a companion paper.) Everyone then sits the same standardised test — a CEFR language score and an analytical score — and then a human recruiter reads the transcript, listens to the audio, sees the test scores, rates the interview 1 to 3 with a written justification, and decides. When a human ran the interview, that same human evaluates it. Recruiters know which arm the applicant came from.

So only one thing is randomised: who asks the questions. The information-collection stage is automated; the evaluation stage is not. This is the design detail that makes everything downstream interpretable, and it is also the paper’s claim to novelty — the existing algorithmic-hiring literature automates the résumé screen before the interview or the recommendation after it, and leaves the conversation alone.

What happened. Applicants interviewed by the machine were 12% more likely to get an offer, 18% more likely to actually start the job, and 18% more likely to still be there thirty days later, all at p < 0.001. The gains persist at sixty, ninety and a hundred and twenty days.

Six small bar charts comparing AI and human interviewers on offers, job starts and retention at 30, 60, 90 and 120 days; the AI bar is taller in every panel, with non-overlapping confidence intervals.
Figure 1, the arXiv draft p. 17: “Treatment effect on key recruiting outcomes in the unconditional sample”

Thirty-day retention is not an academic proxy the authors picked for convenience; it is the metric the clients pay on. The obvious worry is that this is just volume — more offers mechanically produce more warm bodies, of lower average quality. The paper pushes back three ways. Conditional on having accepted an offer, AI-arm applicants are still more likely to start and to last a month. Among those who eventually left, the voluntary-versus-involuntary split is identical across arms, so the AI hires are not being fired more. And on the three KPIs that jointly define productivity in a call centre — average handle time, quality assurance, customer satisfaction — there is nothing: no significant difference, no consistent direction, small magnitudes. “This gain in match quality,” as Jabarian said, “is not at the expense of productivity.”

Controlled variance, which is the mechanism and also the punchline. The transcripts explain how. The AI covers 45% of the 14 topics on average against the humans’ 38%. Its ordering tracks the guideline sequence at τ=0.53\tau = 0.53 versus τ=0.33\tau = 0.33. Its wording sits closer to the sample questions, 0.59 against 0.43. And on every one of these it has significantly lower variance across interviews.

Six histograms in two rows comparing human and AI interviewers on percentage of the 14 topics covered, correlation with the guideline topic order, and similarity to the guideline questions; the AI row is shifted right and more concentrated in every column.
Figure 2, the arXiv draft p. 22: “Recruiter distribution topic coverage”

Notice what these numbers are not. A correlation of 0.53 is not a robot reading a script; it is something that goes off-topic constantly and then comes back. That is the authors’ term of art: controlled variance, the agent adapting inside a standard rather than eliminating adaptation. And the consistency does not come from flattening the language — the AI’s vocabulary richness is higher than the humans’ and more tightly distributed, which is the sort of result that is hard to argue with and mildly insulting to everyone involved. Applicants respond accordingly: they produce more of the linguistic features that predict offers in human-led interviews, like sustained exchange, and fewer of the ones that predict rejection, like backchannel noises and asking questions.

Now the good part. Before any of this was known, the firm surveyed its recruiters and asked them to forecast. They were pessimistic — 61% expected AI-led interviews to be of lower quality. Then they scored them: mean interview rating 1.90 for their own interviews, 2.01 for the machine’s, and the written justifications turned measurably more positive. This is not a handful of enthusiasts; 69% of recruiters make more offers off AI interviews than off their own.

And then, in the actual decision, they quietly discounted the thing they had just praised.

A regression table whose interaction rows show the interview score becoming a weaker predictor of job offers and the language test score a stronger one when the interview was AI-led.
Table 2, the arXiv draft p. 34: “Predicting job offer decisions of recruiters”

Regress the offer decision on the three standardised signals and interact each with the AI arm. The interview score is worth 0.091 in general; interacted with AI it loses 0.047. The language test is worth 0.108; interacted with AI it gains 0.028. So when the machine ran the conversation, the recruiter leans away from the interview and toward the independent test — while simultaneously rating that interview more highly than one they conducted themselves. They compliment the signal and mark it down in the same sitting. The effect is concentrated, encouragingly for the authors’ reading, among the recruiters who say in the survey that interviews matter more than tests. Jabarian’s framing is that the friction is not aversion to AI recommendations — nobody is being told what to do here — but something subtler about the provenance of raw information: “how to train humans to basically read and trust more AI signal.”

The firm got slower. The agent answers the phone at any hour, and it shows: median time from expression of interest to interview falls from 0.51 days to 0.32. Then the interviews pile up in front of the humans who have to evaluate conversations they did not have, and median interview-to-offer time goes from 2.62 days to 7.24. Total median time to hire: 24 days under AI against 20 under humans. This is a queue with a moved bottleneck, and the paper is admirably unembarrassed about it. You automated the fast part.

The cost appendix is similarly deflating in a useful way. Break-even on the one-time deployment is

n=FcHcAI,F=$10,000n^{*} = \frac{F}{c_H - c_{AI}}, \qquad F = \$10{,}000

where nn^{*} is the number of interviews needed to repay the fixed deployment cost FF given the per-interview gap between a human (cHc_H) and the agent (cAIc_{AI}), equation (1) of the arXiv draft. It is undefined in exactly one of the nine calibrated wage-and-price cells — low wages, expensive AI — and the Philippines is the low-wage calibration. Which is the quiet structural point: this deployment was not primarily a labour-cost arbitrage. The gap it exploited was a quality gap, in a place where human interviewers are cheap.

The objections. McElheran, discussing, was warm about the craftsmanship — “this is pre-registered. It’s big. This is really carefully done. Um, we need more of this” — and liked that the paper “spans the full funnel” and digs into mechanism “so that we’re not just sprinkling magic AI fairy dust on an existing workflow.” Her concerns were three. First, generalisation: this is “a narrow setting with kind of a modular point solution,” and “I think we need to be upfront about where this could and could not apply.” Second, she had gone and read the appendix, and found that candidate unavailability is significantly higher in the AI arm — alongside the two failure modes with no human counterpart, the 7% of AI interviews aborted by a technical fault and the 5% ended by applicants who refused to keep talking to a machine. “There’s a bunch of stuff that’s sort of failing in the AI arm and I’m not entirely buying that this is just sort of random or exogenous.”

Third, and best: “I’m wondering how expert this specific interview job is, if reducing variance and homogenizing it is really what we want.” She then improvised past her slides into the thing that actually matters — humans go off topic, and sometimes that is the point. “I know that sometimes when I’m in a conversation with my kids and they go off topic, it’s not so great if I push them back immediately to where they were. And so maybe thinking through the performance metric is potentially a bit narrow.”

Erik Brynjolfsson, given the floor first by the chair, picked the same thread up and made it an economics question: there are productions “where you’re actually trying to maximize variance.” When he looks for PhD students he wants a rock star; the venture capitalists a few miles up the road would very much rather have fat tails. So what happens to the value of a variance-compressing interviewer when the thing you are hiring for is the tail? Jabarian conceded it cleanly — “I’m not saying variance is always bad. We’re saying when it’s bad” — and then, unprompted, named the setting where he would not want this at all: judges. “You don’t want to standardize your context because you can send people in jail for wrong reasons.” It is a strange and slightly wonderful moment, a candidate for best line of the session: the author of the pro-standardisation paper volunteering the place where his own result should be ignored.

So here is where it lands. A machine did an expert human task somewhat better, on the firm’s own money metric, without making the hires worse at their jobs — and it did so not by being clever but by being the only participant in the process capable of following the instructions everyone had agreed to. The firm’s reward for this was a hiring pipeline that runs four days slower, staffed by people who rate the machine’s work highly and then trust it less. Jabarian’s stated motivation was to find the hardest task he could and watch the AI fail at it. It didn’t, which is roughly the least convenient outcome available, and to his credit he seems to have noticed.