Notes on:

Working with AI: Measuring the Applicability of Generative AI to Occupations

Kiran Tomlinson, Sonia Jaffe, Will Wang, Scott Counts & Siddharth Suri
arXiv:2507.07935
7 May 2026
generative AI · labour markets · occupations · measurement
Talk · Paper · doi · Transcript
Made with AI: Opus 5 (reading and writing)

Part of AI and Economic Measurement, Spring 2026

Kiran Tomlinson, Sonia Jaffe, Will Wang, Scott Counts and Siddharth Suri (Microsoft Research), “Working with AI: Measuring the Applicability of Generative AI to Occupations.” Presented at the NBER conference on AI and Economic Measurement, Stanford, 7 May 2026; discussant Kristina McElheran (University of Toronto), who took all three papers in the session together. Written from the arXiv draft, most recently revised 22 December 2025.

Here is a conversation. A user pastes a command-line error message into Bing Copilot. Copilot replies with instructions for editing a configuration file. Now: what job was just done?

The honest answer is that two jobs were done, by two different parties, and they are not the same job. The user was resolving a computer problem. That is an activity that, in the U.S. Department of Labor’s occupational taxonomy, belongs to programmers and system administrators. The AI was advising others on the use of technologies and explaining technical details of products — which belongs to IT and technical support workers. One conversation, two occupations, and neither of them is “the user’s job.”

This is the load-bearing idea in the paper, and it is the sort of thing that seems obvious once stated and then quietly wrecks a decade of measurement. Every prior estimate of AI exposure asks a single question — could a model do this job? — and answers it with a prediction. Frey and Osborne asked machine-learning researchers. Felten, Raj and Seamans aligned benchmark progress to occupational abilities. Eloundou and coauthors used human raters. Tomlinson and his coauthors instead asked what people are already doing, and got two answers per conversation instead of one. As the speaker put it, the literature’s proxies were “alignments between AI benchmarks and occupational tasks or patents,” and “our thinking was to use our observations of what people are using AI for to get more real-world-grounded AI exposure metrics.”

So the two-sided split gives you a cleaner economic object than the usual “exposure.” The user-goal side measures the work AI assists with: you are still doing your job, with a machine at your elbow. The AI-action side measures the work AI performs: the work of some third party you would otherwise have gone to. The helpdesk. The tutor. The copywriter. That is not a metaphor; it is what the classifier is literally scoring.

The machinery. About 200,000 anonymized, privacy-scrubbed U.S. Bing Copilot conversations from 1 January to 30 September 2024 — roughly 100,000 drawn uniformly, plus 100,000 drawn only from conversations that got a thumbs up or down. Each conversation is mapped, on both sides, into O*NET’s 332 intermediate work activities, which are deliberately cross-occupational: “Edit written materials or documents” is done by editors and lawyers and marketers alike. (The competing Anthropic study of Claude conversations assigns each one to a single occupation-specific task, which forecloses exactly this inference. The paper says so, politely.) A GPT-4o pipeline summarizes each side in O*NET’s house style, embeds it, ranks all 332 activities by similarity, and then does binary match/no-match on every one of them. The average conversation comes out with 2.88 user-goal activities and 6.63 AI-action activities.

Then three success measures: real thumbs feedback, an LLM judgement of whether the task was completed, and — the clever one — scope, a six-point rating of what fraction of all the work in that activity, across all occupations, the model could do given only the capability it just demonstrated. Scope is what stops one limerick from certifying the model as a novelist.

Roll it up and you get, for each occupation, equation (1) in the paper:

aiuser  =  jIWAs(i)1 ⁣[fjuser0.0005]cjusersjuserwij a_i^{\text{user}} \;=\; \sum_{j \,\in\, \mathrm{IWAs}(i)} \mathbb{1}\!\left[f_j^{\text{user}} \ge 0.0005\right]\, c_j^{\text{user}}\, s_j^{\text{user}}\, w_{ij}

where aiusera_i^{\text{user}} is occupation ii’s user-goal applicability score, the indicator switches on only if activity jj makes up at least 0.05% of Copilot activity, cjc_j is the completion rate, sjs_j the share of conversations at moderate-or-better scope, and wijw_{ij} the O*NET importance-and-relevance weight of that activity within that job. In English: an occupation’s score is the share of its work — weighted by how much that work matters to the job — that people are demonstrably already bringing to a chatbot, getting finished, and getting finished across a decent chunk of what the activity actually involves. The AI-action score is defined the same way, except its weights exclude tasks a classifier labelled as requiring touching or moving people or objects.

What comes out. Information work, everywhere. The activities that succeed are communicating, teaching, explaining, writing; the ones that fail are image generation and data analysis. Scope correlates with log activity share at 0.64 on the user side and 0.54 on the AI side, which the authors read as users having already found the activities where the model’s reach is broadest. And AI-action success sits consistently below user-goal success: the model can help with a broader slice of your work than it can do outright.

Box plot of AI applicability scores across the 22 SOC major occupation groups, sorted high to low. Computer and Mathematical, Sales and Related, and Office and Administrative Support cluster near 0.3; the groups fade to near-invisible grey toward the right, ending with Healthcare Support below 0.05. Blue diamonds mark groups where a majority of workers do information work.
Slide at 00:33:30: “it is perhaps unsurprising that we see much higher AI applicability score for occupational groups that consist mainly of information work, with computer and mathematical occupations at the very top, and then a lot of physical occupations at the bottom end.”

At the level of the 93 SOC minor groups, Media and Communication Workers top the table at 0.38, then Sales Representatives (Services) at 0.35 and Information and Record Clerks at 0.33; Forest and Conservation Workers come last at 0.03.

Three-column table listing all 93 SOC minor groups with their AI applicability scores, running from Media and Communication Workers at 0.38 down to Forest and Conservation Workers at 0.03. Group names printed in black are majority information work; grey names are not. Asterisks mark U.S. employment size.
Table 1, paper p. 4: “SOC minor groups by AI applicability score.” Score is the employment-weighted average across occupations in the group, averaging the user-goal and AI-action scores.

Individual occupations rank about as you would guess until they don’t: Interpreters and Translators 0.492, Historians 0.462, Writers and Authors 0.454, Services Sales Reps 0.449 — and then, fifth in the entire economy, Computer Numerically Controlled Tool Programmers at 0.419, which reads blue-collar and turns out, in O*NET’s decomposition, to be mostly writing and checking instructions. The single activity contributing most across the top 25 is “Edit written materials or documents.”

Sankey-style diagram. On the left, twenty intermediate work activities headed by “Edit written materials or documents” and “Respond to customer inquiries”; on the right, the 25 highest-scoring occupations from Interpreters and Translators at 0.492 down to Brokerage Clerks at 0.35; coloured ribbons connect each activity to the occupations it contributes to.
Figure 2, paper p. 5: “Top occupations by AI applicability score and their contributing IWAs.” The 25 highest-scoring SOC occupations and the 20 work activities contributing most to those scores.

The two sides really are different. The user-goal and AI-action activity sets are completely disjoint in 40% of conversations, and in 96% there are more activities unique to one side than shared. Per work activity the user brought, the AI performs two. The most lopsided cases are wonderful: “Purchase goods or services” appears 118 times more often as something the AI is assisting with than performing, “Execute financial transactions” 59 times, “Perform athletic activities” 47 times. In the other direction, “Train others on operational or work procedures” runs 18 times more often as an AI action. The speaker’s gloss on the assisted-but-not-performed pile: “these are things where people are getting AI advice for, but the AI is not connecting lines to each other.” A 2024 chatbot will happily talk you through a wire transfer and will not, under any circumstances, make one. He notes this could flip once agents get tool access, which is the most quietly consequential sentence in the talk.

Scatter of the 22 SOC major groups with user-goal applicability score on the horizontal axis and AI-action applicability score on the vertical, with a dashed 45-degree-ish trend line. Blue segments pull Arts/Design/Media and Business & Financial Operations above the line; red segments pull Food Preparation & Serving and Installation, Maintenance & Repair below it. Marginal labels read “AI performs relevant tasks / more likely to delegate subtasks to AI” and “AI assists with relevant tasks / more likely to collaborate with AI on existing tasks.”
Slide at 00:38:55: “examples where we see high user goal and low AI action applicability are some of these very physical occupational groups like food prep and serving and installation, maintenance and repair.”

Now the surprising part. Line these scores up against Eloundou et al.’s human-rated exposure and you get an employment-weighted occupation-level correlation of 0.73 — the new instrument broadly agrees with the old one. Line them up against wages and the agreement collapses. Employment-weighted, the correlation between AI applicability and average occupational wage is 0.13, or 0.19 dropping the top decile. Split it: 0.05 on the user-goal side, 0.26 on the AI-action side. The entire prior literature found AI exposure climbing steeply with pay.

Two panels. Left: a scatter of AI applicability score against average occupational wage on a log axis, with a nearly flat weighted least-squares line labelled r-sub-w equals 0.13; a yellow box highlights a cluster of high-scoring occupations sitting at $30,000 to $50,000. Right: box plots of applicability score by modal education requirement, showing wide overlapping ranges from “less than high school” to “master’s or higher”. Inset at bottom left, a reproduction of the steeply rising Eloundou et al. exposure-by-income curve.
Slide at 00:36:40: “we see a fairly weak correlation both with average wage and with education requirements, especially after weighting by employment at the occupation level. And this is not something that previous work has always done.”

It isn’t noise. It’s composition, and the composition is enormous. Customer Service Representatives: 2.86 million workers, score 0.408. Services Sales Reps: 1.14 million, 0.449. Office and Administrative Support as a whole: 18.16 million workers, group score 0.26, against Healthcare Support’s 7.06 million at 0.05. “This is driven by some low wage, low education requirement occupations that are very large and that have high AI applicability,” the speaker said, “and these are clustered in sales and office and admin support.”

And then the knife twist. Drop the employment weights and the correlation climbs back to 0.17 on the user side and 0.32 on the AI side — roughly the old finding. Which is to say a chunk of “AI hits high-wage work” may be an artifact of counting occupations rather than workers. Why does that matter so much? Because occupation boundaries are a bureaucratic accident. Cooks are split into Fast Food, Restaurant, Institution and Cafeteria, and Short Order — four data points. Maids and Housekeeping Cleaners are one. An unweighted regression treats those five as five equal observations and does not notice that one of them contains 836,000 people. A finding that reads like an economic law turns out to be partly a fact about how the Standard Occupational Classification chops things up.

The paper is unusually honest, which is worth noticing given who wrote it. These are Microsoft Research employees measuring Microsoft’s product; the aggregates are public and the disclaimers are more prominent than the results. Three of the authors hand-labelled 195 conversations and report that they agreed with each other at Cohen’s κ of 0.41 to 0.58, and with the pipeline at 0.34 to 0.53 — moderate, and reported as such. Scope agreement is worse: human–human within-group agreement ranges 0.12 to 0.66, human–LLM 0 to 0.49, with mean absolute error around 1.1 on a six-point scale where random guessing gives 1.94. And in the single best act of preemptive self-sabotage in the genre, the authors show that by moving the usage threshold from 1% of chat activity to 0.01%, you can conclude that either ~0% or ~100% of the workforce has half its importance-weighted tasks represented in the data. The rankings are stable across three orders of magnitude of threshold. The headline percentage is not, so they decline to produce one. Onstage, unprompted, the speaker flagged that the measures “are not indicative of what occupations AI is likely to replace” and name-checked a Washington Post article about the paper as the reason he was saying so. (The paper’s version of the argument is the ATM: automating the bank teller’s core task raised the number of teller jobs, because branches got cheaper to open.)

The discussion, and its absence. McElheran discussed all three session papers in one nineteen-minute block, and her theme was the level of analysis. “One of the things that you have to start looking for when you’re consuming all of this literature is where the word occupation, where the word tasks, and where the word jobs come in, because they don’t always mean the same thing.” Her exhibit was Frey and Osborne’s 47%-of-jobs-at-risk versus Arntz, Gregory and Zierahn, who ran the same machinery on survey data about what workers actually do in their jobs and watched the number fall from a baseline of 38% to 9%. Her verdict on the state of the field: “we have built this edifice of thinking about exposure and worker implications that is all on this one narrow and I’m going to use Claude’s favorite term load-bearing column of O*NET data at the occupation level.” Aimed here specifically: “there’s a lot of places where the sleight of hand happens and we think we’re talking about workers, but we’re actually talking about tasks in occupations.”

On this paper’s signature move she was warm and unconvinced: “I think this user goal versus AI action is distinctive and conceptually novel, but I wasn’t entirely sure what to do with it.” Her most actionable ask was statistical — “some of these things are within the confidence interval and I think you probably want to do some bootstrapping” — and she is right that no ranking in the paper carries an interval, which sits awkwardly beside a scope classifier whose human agreement bottoms out at zero. She also wanted the wage mechanism rather than the coefficient. And she delivered, as a joke, the sharpest thing anyone said all session: “now we’re going to have AI classifying O*NET stuff so we can use the AI-generated O*NET data to study AI. That’s going to be an interesting circular thing here.” The paper’s own footnote reports “strong evidence that GPT-4o includes O*NET data in its pretraining corpus.” The classifier has read the taxonomy it is being graded against. The authors treat this as a convenience.

Here is the honest part about the exchange: there wasn’t one. Some of these points already have standing answers — the applicability-is-not-replacement disclaimer came about forty-five minutes before her remarks, and Appendix G lays out the employment-weighting mechanism she asked for — but the authors never responded on the record. The general Q&A went almost entirely to the other two papers in the session; the chair explicitly asked for questions on the Microsoft paper and the Bick adoption survey, and the next question went to the AI-interviewer field experiment anyway. The one moment that brushed this paper came from a questioner the chair seemed to call Harry (the caption is unreliable and no surname was given), who said community colleges training people for health care and IT jobs “are terrified about all this. All of them are. So given the state of the knowledge, given the state of the data, what’s the best thing we can tell them?” McElheran took it, not the authors, and answered from her own work on older workers displaced by earlier technology waves: a bifurcation is coming between people comfortable with the tools and people who are “scared, they don’t understand it, they don’t like it, they don’t like how it was trained,” and the latter are at real risk of being left behind.

The talk also carried one result not in the December draft: recomputing everything month by month across the nine-month window shows the entropy of the activity distribution rising through 2024, with applicability scores rising alongside it — people spreading their usage out rather than concentrating it. The slide said “Over 6 months.” The speaker corrected it live: “this is a typo. This should say nine months.”

Which leaves the thing everyone in that room already half-knew. This paper’s most quotable finding — the wage correlation that vanishes — is a finding about weighting. Its most careful contribution is a refusal to publish the number that would have gotten it cited most. And the instrument doing the measuring learned the taxonomy from the same internet it is now being asked to score.