In April 2023, Erik Brynjolfsson, Danielle Li and Lindsey R. Raymond published Generative AI at Work as NBER Working Paper 31161. Their claim was specific: an AI assistant increased customer support agents’ productivity by 14% on average. The improvement reached 34% for novice and low-skilled workers, with minimal effects on experienced, highly skilled workers.
The record supports that unequal distribution of gains. It does not support treating the average improvement as every worker’s result, or treating faster resolution as proof that experience no longer matters.
What counted as performance?
The researchers studied a staggered deployment at a company supplying business software. Its support agents handled customer problems through online chats. The assistant supplied suggested responses and information during those conversations. Agents remained responsible for the replies and could disregard its suggestions.
The principal productivity measure was issues resolved per hour. That combines the pace of handling conversations with whether agents resolve the problems. It is not a count of words generated or suggestions accepted.
The April 2023 manuscript covered 5,179 agents and reported the 14% average gain. The subsequently published Quarterly Journal of Economics article used 5,172 agents and reported 15%. Those are different versions of the analysis, not interchangeable figures. The experience-related pattern survived publication: less-experienced and lower-skilled workers gained more.
That distinction matters. A department-wide average can conceal an intervention that mostly improves the performance of workers who previously needed more time.
Who improved most?
The researchers examined both tenure and prior performance. These are different attributes. A recently hired agent lacks time in the job. A lower-performing agent has weaker measured results. Neither label establishes that the worker lacks general ability.
In the April 2023 analysis, the largest gains accrued to newer and lower-skilled agents. The widely cited 34% figure belongs to that finding. It should not become a promise that every new hire will improve by that amount.
The tenure comparison was particularly concrete. The authors reported that assisted agents with 2 months of tenure performed comparably to unassisted agents with more than 6 months of tenure. The assistant compressed part of the observed learning curve.
Experienced, high-performing agents had much less room to improve. Their small gains are essential to the conclusion, not an inconvenient exception. If all groups had improved equally, the study would show higher productivity without necessarily showing a narrower performance gap.
Here, the less-experienced workers moved closer to the output levels previously associated with longer tenure. That is the evidence behind the claim.
How strong was the comparison?
This was a workplace deployment study, not a trial randomly assigning individual agents to receive the assistant. The researchers used differences in rollout timing to compare performance before and after access, alongside agents who had not yet received it.
Their analysis accounted for persistent differences between workers and common changes over time. They also examined performance around the introduction of the tool. This is substantially more informative than comparing an assisted department’s results with its previous monthly total.
The causal interpretation nevertheless depends on the comparison capturing what would have happened without access. An unmeasured change arriving with the rollout could complicate that interpretation. The study also concerns one company and one support setting. Its estimates are not a representative measurement of every occupation using generative AI.
Within that design, the experience-level comparisons identify who gained rather than merely documenting an increase in the company’s aggregate output.
Did speed come at quality’s expense?
A support agent can finish quickly by giving an inadequate answer. The paper therefore examined more than handling time. Resolution measures and customer responses helped distinguish useful acceleration from conversations ending sooner.
The published article reports that less-experienced and lower-skilled workers improved both speed and quality. For the most experienced and highest-skilled workers, it reports small speed gains alongside small quality declines.
That result rules out a simple description of the assistant as an equal improvement for everyone. Assistance could help a novice reach a better answer while offering an expert a response below the expert’s usual standard.
The researchers also found improvements in customer sentiment and fewer requests to speak to a supervisor. Those outcomes matter because the customers, not just the employer’s throughput measure, registered changes in the interaction. The quality findings strengthen the case for novice gains while weakening any blanket claim about expert performance.
Was the assistant transferring expertise?
The authors’ proposed mechanism was the spread of useful practices from stronger workers. The system drew on historical support conversations and made suggestions available while an agent was handling a problem.
That arrangement changes when knowledge becomes accessible. An agent need not first acquire every useful response through months of encounters. The assistant can supply relevant language or information during the encounter itself.
The researchers examined changes in conversational language and performance during periods when the system was unavailable. They found evidence consistent with workers learning from assistance rather than only borrowing its output moment by moment.
That is suggestive evidence of knowledge transfer, not a direct measurement of everything an agent understood or retained. The defensible finding is that access accelerated demonstrated performance, with some evidence that the benefit persisted beyond immediate suggestions.
Does other research show the same pattern?
Related experiments help separate this result from a universal rule. In their 2023 Science study of professional writing tasks, Shakked Noy and Whitney Zhang found that ChatGPT improved productivity and reduced inequality between participants. Weaker initial performers benefited more. That resembles the support study’s distributional finding, although the tasks and measures differed.
Sida Peng and coauthors tested GitHub Copilot through a programming task. Participants with access completed it faster. That supplies experimental evidence for assistance improving task speed, not a replication of customer support tenure effects.
Anil Doshi and Oliver Hauser’s short-story experiment found that generative AI assistance particularly benefited less-creative writers while making the resulting stories more similar to one another. Better individual scores did not imply improvement on every dimension.
Michelle Vaccaro, Abdullah Almaatouq and Thomas Malone’s meta-analysis likewise found that human-AI combinations did not reliably outperform whichever was better alone, the human or the AI. Together, these studies support testing the worker, task and outcome separately. They do not supply a universal novice bonus.
Does a smaller gap mean fewer jobs?
The support study measured performance under assistance. Issues resolved per hour cannot, by itself, establish what employers will do with staffing, wages or the additional capacity.
Tyna Eloundou and coauthors’ GPTs are GPTs assessed occupational exposure to language models. Exposure describes tasks that technology might affect. It is not an observed employment loss. Substituting an exposure estimate for this deployment’s measured productivity gain would mix different questions.
David Autor’s analysis of workplace automation explains why automating tasks can also increase the value of complementary human work. Daron Acemoglu’s The Simple Macroeconomics of AI makes the aggregation problem explicit: economy-wide gains depend on which tasks change and how much their costs fall.
The support findings establish neither a staffing cut nor a wage increase. A narrower output gap does not determine who receives the financial benefit.
What is the verdict?
Yes. The AI assistant narrowed the measured performance gap because newer and lower-performing support agents improved substantially more than experienced, high-performing colleagues. The strongest evidence is not the average productivity increase. It is the steeper improvement among workers still acquiring the job’s practices. The assistant made some expertise easier to use. It did not make expertise unnecessary.



