Back to Research News
Editorial cover for a higher-education experiment with GPT-4 feedback
Journal PaperPeer-reviewed study202617 Jul 2026· 3 min

GPT-4 feedback increases student activation and learning outcomes in higher education

Stephan Geschwind, Johann Graf Lambsdorff, Deborah Voss, Veronika Hackl

International Journal of Artificial Intelligence in Education

500-word summary

Editorial cover for a higher-education experiment with GPT-4 feedback

Listen to the 500-word paper summary

Audio summary

0:00/0:00

Geschwind, Lambsdorff, Voss, and Hackl investigate whether generative AI can make individualized feedback more scalable in large university courses. Their 2026 open-access IJAIED article reports a lab-in-the-field experiment conducted in undergraduate macroeconomics tutorial classes over one semester. Students answered eight open-ended questions and experienced one of three feedback arrangements: classroom-level lecturer feedback, additional individual feedback from peers, or additional individual feedback generated by GPT-4. The study asks whether the feedback changes student activation and the quality of subsequent answers, rather than only whether students say they like AI.

All groups received lecturer discussion of a sample solution and adaptive feedback on two selected student answers. In the peer-feedback condition, students anonymously rated another student's answer for content and style and wrote suggestions for improvement. In the AI-feedback condition, GPT-4 provided individual scores and qualitative suggestions in a comparable structure. Both individual-feedback formats were designed around three familiar feedback functions: looking back at current performance, clarifying the goal through a sample solution, and pointing forward to improvement.

The researchers operationalized activation in two ways. First, they tracked voluntary participation across the eight tasks; tutorial participation did not count toward the final examination, although students who completed at least seven tasks entered a raffle. Second, they used answer length as an indicator of the effort invested in a task. For learning outcomes, GPT-4 rated the content and style of answers on five-point scales. To make treatment comparisons more consistent, the model rated answers from all three conditions after the course in three separate iterations, and the researchers averaged those ratings.

The AI-feedback condition produced the clearest activation pattern. Participation remained higher over time than in the lecturer-only condition, while peer feedback performed only marginally better than lecturer feedback. Students receiving AI feedback also wrote the longest answers. Because attrition was voluntary and not random, the authors did not rely only on simple averages: their task-to-task analysis used 479 repeated comparisons in which the same student participated in the same condition on consecutive relevant tasks.

For learning outcomes, AI feedback was associated with the strongest improvement in content ratings. Peer feedback did not produce the same advantage, which the authors connect to differences in reliability and completeness: peers sometimes failed to provide feedback, while the AI returned it consistently. The study found no comparable treatment effect on writing style. That distinction matters because it suggests that the intervention supported substantive engagement with macroeconomic answers without automatically improving every dimension of academic writing.

The paper is promising but should be interpreted carefully. The setting was one subject area and one university course, participation was voluntary, attrition differed across conditions, and the outcome measure was improvement in open-ended answers rather than a broad examination of long-term retention or transfer. GPT-4 generated feedback in one condition and also served as the common post-course rater, even though the authors used repeated ratings and prior reliability work to strengthen the measure. These choices make the study more informative than a satisfaction survey, but they do not establish that AI feedback will outperform expert human feedback in every context.

For AIEDHK, the practical lesson is to focus on feedback system design. Individual AI feedback may sustain participation when it is timely, structured, connected to a sample solution, and followed by opportunities to revise. Educators should preserve lecturer oversight, test feedback validity, measure learning with independent assessments, and compare the intervention with realistic alternatives. The contribution is not a claim that GPT-4 replaces teachers. It is evidence that carefully structured AI feedback can extend the reach of tutorial support while keeping course goals and human judgment visible.

Related papers

Four diverse university students practise prompting and source checking with an instructor at a library learning table
Journal Paper2026
Journal Paper 52

A 90-minute GenAI literacy course improved knowledge, prompting, source checking and self-efficacy across 65 university sections

Allison E. Connell Pensky, Lydia E. Eckstein, Michael C. Melville, Laura O. Pottmeyer, Zach Mineroff, Avi Chawla, Judy Brooks, Chad Hershock, Marsha C. Lovett

Computers & Education

In a large experiment involving 1,368 undergraduate and graduate students across 65 university course sections, a 90-minute asynchronous GenAI learning module improved knowledge of how the technology works, prompt-engineering performance, fact- and source-checking, and self-efficacy. It did not improve critical evaluation of bias, showing that short foundational training needs deeper practice for responsible judgment.

generative AI literacyrandomized experimenthigher education
Read 500-word summary
A university student compares an AI explanation with handwritten concept notes while an instructor and peers work in a seminar room
Journal Paper2026
Journal Paper 50

Experimental evidence on the learning impact of generative AI: gains persisted when students used it for explanation rather than automation

Zara Contractor, Germán Reyes

arXiv working paper

A randomized, proctored experiment reported that undergraduate access to off-the-shelf generative AI raised immediate factual and conceptual test performance by 0.27 standard deviations and that the gains persisted one week later. The working paper also finds a consequential usage pattern: students who used AI to explain concepts showed stronger delayed gains than students who used it to automate drafting.

generative AIrandomized experimenthigher education
Read 500-word summary
Editorial cover of undergraduate learners and a lecturer examining a course-grounded RAG chatbot alongside flat learning and motivation outcome traces
Journal Paper2026
Journal Paper 36

AI chatbots in higher education: Comparing expectations to evidence

Andrew Thoeni, Luke K. Fryer

Computers in Human Behavior Reports

A semester-long randomized field experiment with 454 undergraduates found that access to a course-grounded RAG chatbot did not significantly improve interest, self-efficacy, engagement, or test performance, despite students reporting that they liked the tool.

RAG chatbotrandomized field experimenthigher education
Read 500-word summary