返回研究新聞
Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
期刊論文同行評審研究20262026年8月9日· 9 min

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

500 字摘要

Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory

Cruces and colleagues ask whether generative AI widens or narrows performance differences associated with formal education. Their August 2026 arXiv working paper reports a preregistered online experiment with 1,174 adults aged 25 to 45 in Argentina. Participants were classified into lower- and higher-education groups using a preregistered threshold, then randomly assigned to complete an incentivized workplace-style business problem with or without an embedded GPT-4.1 assistant. Everyone then completed an immediate follow-up module without AI. This design distinguishes assisted task performance from what participants could articulate or recall once assistance was removed.

The task required participants to read an email from a hypothetical manager, examine text, a figure and a table, diagnose a problem and propose a solution. It was self-contained and designed to draw on reading, data comprehension, reasoning, creative problem-solving and writing rather than specialized industry knowledge. In the control condition, higher-education participants outperformed lower-education participants by 0.548 standard deviations. With AI, the gap fell to 0.139 standard deviations, a reduction of about 75 percent. Relative to the lower-education control group, AI raised the overall task score by 1.242 standard deviations for lower-education participants and 0.834 for higher-education participants.

The gap did not disappear. Chat-log analysis suggests that lower-education participants obtained substantial assistance, while higher-education participants used the assistant somewhat more effectively through more detailed prompts and more structured workflows. The follow-up also complicates a simple delegation explanation. Treated participants did not perform worse after AI was removed. Lower-education participants retained a modest 0.171-standard-deviation gain, while the higher-education estimate was small and not statistically significant. Yet a 0.200-standard-deviation education gap re-emerged in the unassisted follow-up.

Effort is the educationally important mechanism. Intensive AI assistance predicted strong performance on the main task even when participants invested less of their own effort. Better unassisted follow-up performance, however, appeared when intensive assistance was combined with sustained task engagement. The study therefore separates effective performance from underlying human capital: AI can lower the expertise needed to complete a task, but durable understanding still depends on reading, evaluating, integrating and explaining information.

The evidence is bounded. This is a working paper, not yet a peer-reviewed journal article. The experiment lasted about 21 minutes on average, used one business scenario, measured an immediate follow-up and recruited adults rather than school students. It does not establish long-term skill growth, labor-market outcomes or effects in other languages and institutions. The education-group threshold reflects the Argentine context, and the comparison does not show that formal education itself caused every observed difference in strategy or performance.

For Hong Kong education and workforce training, pilots should report both assisted and independent outcomes. Learners can use AI to compare sources and draft a diagnosis, then complete an unaided explanation or transfer task. Access may reduce short-run gaps, but equitable learning requires supports that help every participant build the prompt, reasoning and verification practices that remain useful when the tool is gone.

相關論文

Four diverse university students practise prompting and source checking with an instructor at a library learning table
期刊論文2026
期刊論文 52

A 90-minute GenAI literacy course improved knowledge, prompting, source checking and self-efficacy across 65 university sections

Allison E. Connell Pensky, Lydia E. Eckstein, Michael C. Melville, Laura O. Pottmeyer, Zach Mineroff, Avi Chawla, Judy Brooks, Chad Hershock, Marsha C. Lovett

Computers & Education

In a large experiment involving 1,368 undergraduate and graduate students across 65 university course sections, a 90-minute asynchronous GenAI learning module improved knowledge of how the technology works, prompt-engineering performance, fact- and source-checking, and self-efficacy. It did not improve critical evaluation of bias, showing that short foundational training needs deeper practice for responsible judgment.

generative AI literacyrandomized experimenthigher education
閱讀 500 字摘要
A university student compares an AI explanation with handwritten concept notes while an instructor and peers work in a seminar room
期刊論文2026
期刊論文 50

Experimental evidence on the learning impact of generative AI: gains persisted when students used it for explanation rather than automation

Zara Contractor, Germán Reyes

arXiv working paper

A randomized, proctored experiment reported that undergraduate access to off-the-shelf generative AI raised immediate factual and conceptual test performance by 0.27 standard deviations and that the gains persisted one week later. The working paper also finds a consequential usage pattern: students who used AI to explain concepts showed stronger delayed gains than students who used it to automate drafting.

generative AIrandomized experimenthigher education
閱讀 500 字摘要
Chinese secondary students complete homework with digital assistance before taking a separate closed-book assessment observed by a teacher
期刊論文2026
期刊論文 68

Generative AI adoption was linked to higher homework scores but lower unaided exams in a 26,811-student panel

David Strömberg, Victor Lei, Yanhui Wu

CEPR Discussion Paper No. 21577

A CEPR discussion paper analyzes 30 months of records from 26,811 Chinese students in Grades 7–12. Its difference-in-differences estimates associate generative-AI adoption with homework scores 18% higher and completion time 30% lower, but with substantial declines on closed-book and entrance examinations.

generative AIsecondary educationhomework outsourcing
閱讀 500 字摘要