← Back to Research News
Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
Journal PaperPeer-reviewed study20269 Aug 2026· 2 min

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

500-word summary

Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory

Listen to the paper summary

Audio summary

0:00/0:00

Cruces and colleagues ask whether generative AI widens or narrows performance differences associated with formal education. Their August 2026 arXiv working paper reports a preregistered online experiment with 1,174 adults aged 25 to 45 in Argentina. Participants were classified into lower- and higher-education groups using a preregistered threshold, then randomly assigned to complete an incentivized workplace-style business problem with or without an embedded GPT-4.1 assistant. Everyone then completed an immediate follow-up module without AI. This design distinguishes assisted task performance from what participants could articulate or recall once assistance was removed.

The task required participants to read an email from a hypothetical manager, examine text, a figure and a table, diagnose a problem and propose a solution. It was self-contained and designed to draw on reading, data comprehension, reasoning, creative problem-solving and writing rather than specialized industry knowledge. In the control condition, higher-education participants outperformed lower-education participants by 0.548 standard deviations. With AI, the gap fell to 0.139 standard deviations, a reduction of about 75 percent. Relative to the lower-education control group, AI raised the overall task score by 1.242 standard deviations for lower-education participants and 0.834 for higher-education participants.

The gap did not disappear. Chat-log analysis suggests that lower-education participants obtained substantial assistance, while higher-education participants used the assistant somewhat more effectively through more detailed prompts and more structured workflows. The follow-up also complicates a simple delegation explanation. Treated participants did not perform worse after AI was removed. Lower-education participants retained a modest 0.171-standard-deviation gain, while the higher-education estimate was small and not statistically significant. Yet a 0.200-standard-deviation education gap re-emerged in the unassisted follow-up.

Effort is the educationally important mechanism. Intensive AI assistance predicted strong performance on the main task even when participants invested less of their own effort. Better unassisted follow-up performance, however, appeared when intensive assistance was combined with sustained task engagement. The study therefore separates effective performance from underlying human capital: AI can lower the expertise needed to complete a task, but durable understanding still depends on reading, evaluating, integrating and explaining information.

The evidence is bounded. This is a working paper, not yet a peer-reviewed journal article. The experiment lasted about 21 minutes on average, used one business scenario, measured an immediate follow-up and recruited adults rather than school students. It does not establish long-term skill growth, labor-market outcomes or effects in other languages and institutions. The education-group threshold reflects the Argentine context, and the comparison does not show that formal education itself caused every observed difference in strategy or performance.

For Hong Kong education and workforce training, pilots should report both assisted and independent outcomes. Learners can use AI to compare sources and draft a diagnosis, then complete an unaided explanation or transfer task. Access may reduce short-run gaps, but equitable learning requires supports that help every participant build the prompt, reasoning and verification practices that remain useful when the tool is gone.

Related papers

A programming lecturer and two diverse university students inspect compiled code, an inheritance diagram, and a grading rubric in a computer laboratory
Journal Paper2026
Journal Paper 102

Five AI systems outscored the average OOP cohort but still failed compilation and advanced concepts

Marina Lepp, Joosep Kaimre

arXiv preprint

Lepp and Kaimre evaluated ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and Microsoft 365 Copilot on authentic introductory OOP tests and examinations using student grading criteria. Systems exceeded the historical average and often solved long tasks, yet some code did not compile and interfaces, abstract classes, inheritance, and image-based questions remained difficult. The results challenge take-home assessment validity without proving student learning.

programming assessmentobject-oriented programminggenerative AI
Read 500-word summary →
A diverse group of university students explores a branching media-technology learning story while an instructor traces where quiz choices connect to the narrative
Conference Paper2026
Conference Paper 90

AI-generated learning stories were clear and well paced, but their quizzes did not belong in the plot

Finn Rogosch, Andreas Schrader

EDULEARN26 Proceedings

Rogosch and Schrader tested AI-generated interactive-fiction episodes with 22 STEM higher-education participants. The five-to-ten-minute stories were rated clear and appropriately long, but story-content coherence averaged below the neutral midpoint and engagement sat near it. Participants most often questioned why characters suddenly demanded technical answers, showing that a playable educational story can still fail to integrate its learning task.

interactive fictioneducational gamesgenerative AI
Read 500-word summary →
A university student compares an AI explanation with handwritten concept notes while an instructor and peers work in a seminar room
Journal Paper2026
Journal Paper 50

Experimental evidence on the learning impact of generative AI: gains persisted when students used it for explanation rather than automation

Zara Contractor, Germán Reyes

arXiv working paper

A randomized, proctored experiment reported that undergraduate access to off-the-shelf generative AI raised immediate factual and conceptual test performance by 0.27 standard deviations and that the gains persisted one week later. The working paper also finds a consequential usage pattern: students who used AI to explain concepts showed stronger delayed gains than students who used it to automate drafting.

generative AIrandomized experimenthigher education
Read 500-word summary →