
Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward
Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi
arXiv working paper
Resumo de 500 palavras

Cruces and colleagues ask whether generative AI widens or narrows performance differences associated with formal education. Their August 2026 arXiv working paper reports a preregistered online experiment with 1,174 adults aged 25 to 45 in Argentina. Participants were classified into lower- and higher-education groups using a preregistered threshold, then randomly assigned to complete an incentivized workplace-style business problem with or without an embedded GPT-4.1 assistant. Everyone then completed an immediate follow-up module without AI. This design distinguishes assisted task performance from what participants could articulate or recall once assistance was removed.
The task required participants to read an email from a hypothetical manager, examine text, a figure and a table, diagnose a problem and propose a solution. It was self-contained and designed to draw on reading, data comprehension, reasoning, creative problem-solving and writing rather than specialized industry knowledge. In the control condition, higher-education participants outperformed lower-education participants by 0.548 standard deviations. With AI, the gap fell to 0.139 standard deviations, a reduction of about 75 percent. Relative to the lower-education control group, AI raised the overall task score by 1.242 standard deviations for lower-education participants and 0.834 for higher-education participants.
The gap did not disappear. Chat-log analysis suggests that lower-education participants obtained substantial assistance, while higher-education participants used the assistant somewhat more effectively through more detailed prompts and more structured workflows. The follow-up also complicates a simple delegation explanation. Treated participants did not perform worse after AI was removed. Lower-education participants retained a modest 0.171-standard-deviation gain, while the higher-education estimate was small and not statistically significant. Yet a 0.200-standard-deviation education gap re-emerged in the unassisted follow-up.
Effort is the educationally important mechanism. Intensive AI assistance predicted strong performance on the main task even when participants invested less of their own effort. Better unassisted follow-up performance, however, appeared when intensive assistance was combined with sustained task engagement. The study therefore separates effective performance from underlying human capital: AI can lower the expertise needed to complete a task, but durable understanding still depends on reading, evaluating, integrating and explaining information.
The evidence is bounded. This is a working paper, not yet a peer-reviewed journal article. The experiment lasted about 21 minutes on average, used one business scenario, measured an immediate follow-up and recruited adults rather than school students. It does not establish long-term skill growth, labor-market outcomes or effects in other languages and institutions. The education-group threshold reflects the Argentine context, and the comparison does not show that formal education itself caused every observed difference in strategy or performance.
For Hong Kong education and workforce training, pilots should report both assisted and independent outcomes. Learners can use AI to compare sources and draft a diagnosis, then complete an unaided explanation or transfer task. Access may reduce short-run gaps, but equitable learning requires supports that help every participant build the prompt, reasoning and verification practices that remain useful when the tool is gone.


