Back to Research News
Editorial cover of undergraduate learners and a lecturer examining a course-grounded RAG chatbot alongside flat learning and motivation outcome traces
Journal PaperPeer-reviewed study202625 Jul 2026· 4 min

AI chatbots in higher education: Comparing expectations to evidence

500-word summary

Editorial cover of undergraduate learners and a lecturer examining a course-grounded RAG chatbot alongside flat learning and motivation outcome traces

Listen to the 500-word paper summary

Audio summary

0:00/0:00

Thoeni and Fryer test a claim that universities increasingly encounter in product demonstrations: if a generative-AI chatbot is grounded in trusted course content, available whenever students need it, and capable of answering questions or generating practice, will it improve learning? Their 2026 open-access article in Computers in Human Behavior Reports reports a semester-long randomized field experiment rather than a satisfaction survey. The result is important precisely because it is negative: the retrieval-augmented generation, or RAG, chatbot did not produce a statistically significant improvement in any of the measured learning-related outcomes.

The experiment took place in three sections of an introductory Principles of Marketing course at a state university in the southeastern United States. One section was asynchronous online, one was a small face-to-face class, and one was a large face-to-face class. After excluding students who dropped or were repeating the course, the study included 454 undergraduates: 231 in the control group and 223 in the treatment group. Students were randomly assigned within each class section after the add/drop period, helping balance the conditions across delivery mode and class size.

The design covered a 16-week semester, with the chatbot intervention operating for 12 calendar weeks. Before treatment, all students completed the same course work and first test. The control group then continued receiving participation credit for study sessions using custom Quizlet flash cards. The treatment group received equivalent credit for at least one chatbot session per chapter, while retaining access to Quizlet without additional credit. Students in both groups could use their assigned support as often as they wished. This made the comparison closer to adding a course chatbot to realistic study options than to replacing instruction.

The chatbot was more carefully constructed than a general-purpose chat window. It ran in a secure university Microsoft Azure environment using Copilot with GPT-4o. The RAG content included instructor-provided concepts, glossaries, learning objectives, and lecture transcripts organized by chapter, but not test questions or textbook content. System instructions gave the tutor goals, a conversational role, and functions for discussion, explanation, multilingual interaction, and short quizzes. The researchers tested its consistency and accuracy and reported that it answered 199 of 200 course test questions correctly, even though those questions were not included in its knowledge base.

The outcomes came from several sources. Students completed pre- and post-treatment measures of individual interest and self-efficacy. Engagement included emotional, participation, performance, and skill scales, weekly self-reports, electronic-book usage, and chatbot session counts. Academic achievement was represented by standardized scores from a common first test before treatment and a common fourth test at the end of the term. The analyses used difference-in-differences models and controlled for gender, age, race, weekly job hours, and whether the course was face-to-face or asynchronous online.

Across interest, self-efficacy, engagement, and test scores, the critical treatment-by-time interactions were not statistically significant. Self-efficacy rose slightly over the semester for students as a whole, and several engagement measures fell, but neither pattern was attributable to chatbot access. For achievement, the treatment-by-time interaction was also non-significant. In other words, the study did not find that adding the course-grounded chatbot changed the measured trajectory relative to the control condition.

Students nevertheless viewed the tool positively. Treatment students reported high enjoyment, perceived help with course material, and interest in having a similar chatbot available in other courses. Their own ratings were more cautious about whether the chatbot improved grades or interest in marketing. This gap is one of the paper's most useful findings: liking an AI tutor, finding it convenient, or wanting continued access does not establish that it improves learning. Adoption and educational effectiveness must be evaluated separately.

The null result also has boundaries. The study involved one introductory subject, one instructor, and three sections at one university. The control group had a legitimate study aid, which sets a more demanding comparison than no support. The chatbot did not remember prior sessions, limiting personalization and social continuity. The researchers could not fully capture how students changed their broader study habits, and the final test covered different course material from the baseline test even though scores were standardized. More advanced learners, other subjects, alternative tutoring strategies, or stronger integration with classroom activity could produce different results.

For AIEDHK, the practical lesson is to evaluate a designed intervention before scaling a platform contract. A course-grounded chatbot may be accurate and popular yet still add no measurable benefit to existing study support. Pilots should define the expected mechanism, compare against a credible alternative, measure independent achievement and engagement over time, inspect actual usage and substitution effects, and test whether memory or personalization helps without creating unacceptable privacy costs. The paper does not show that RAG tutors can never work. It shows that grounding and availability alone are insufficient evidence of learning value.

Related papers

Four diverse university students practise prompting and source checking with an instructor at a library learning table
Journal Paper2026
Journal Paper 52

A 90-minute GenAI literacy course improved knowledge, prompting, source checking and self-efficacy across 65 university sections

Allison E. Connell Pensky, Lydia E. Eckstein, Michael C. Melville, Laura O. Pottmeyer, Zach Mineroff, Avi Chawla, Judy Brooks, Chad Hershock, Marsha C. Lovett

Computers & Education

In a large experiment involving 1,368 undergraduate and graduate students across 65 university course sections, a 90-minute asynchronous GenAI learning module improved knowledge of how the technology works, prompt-engineering performance, fact- and source-checking, and self-efficacy. It did not improve critical evaluation of bias, showing that short foundational training needs deeper practice for responsible judgment.

generative AI literacyrandomized experimenthigher education
Read 500-word summary
A university student compares an AI explanation with handwritten concept notes while an instructor and peers work in a seminar room
Journal Paper2026
Journal Paper 50

Experimental evidence on the learning impact of generative AI: gains persisted when students used it for explanation rather than automation

Zara Contractor, Germán Reyes

arXiv working paper

A randomized, proctored experiment reported that undergraduate access to off-the-shelf generative AI raised immediate factual and conceptual test performance by 0.27 standard deviations and that the gains persisted one week later. The working paper also finds a consequential usage pattern: students who used AI to explain concepts showed stronger delayed gains than students who used it to automate drafting.

generative AIrandomized experimenthigher education
Read 500-word summary
University students discuss transparent and responsible generative-AI use during a collaborative assignment while a teacher facilitates peer reflection
Journal Paper2026
Journal Paper 84

Perceived classmate GenAI use was associated with lower trust, while perceived AI literacy attenuated the direct link

Zhen Zhang, Jiaying Geng, Chunhui Qi

Behavioral Sciences

A cross-sectional survey of 406 students at two institutions found that perceiving a classmate as using more GenAI was associated with lower perceived warmth, competence and interpersonal trust. Perceived target AI literacy weakened only the direct association, while the design cannot establish that AI use caused distrust.

generative AIinterpersonal trustAI literacy
Read 500-word summary