
AI chatbots in higher education: Comparing expectations to evidence
Andrew Thoeni, Luke K. Fryer
Computers in Human Behavior Reports
Resumo de 500 palavras

Thoeni and Fryer test a claim that universities increasingly encounter in product demonstrations: if a generative-AI chatbot is grounded in trusted course content, available whenever students need it, and capable of answering questions or generating practice, will it improve learning? Their 2026 open-access article in Computers in Human Behavior Reports reports a semester-long randomized field experiment rather than a satisfaction survey. The result is important precisely because it is negative: the retrieval-augmented generation, or RAG, chatbot did not produce a statistically significant improvement in any of the measured learning-related outcomes.
The experiment took place in three sections of an introductory Principles of Marketing course at a state university in the southeastern United States. One section was asynchronous online, one was a small face-to-face class, and one was a large face-to-face class. After excluding students who dropped or were repeating the course, the study included 454 undergraduates: 231 in the control group and 223 in the treatment group. Students were randomly assigned within each class section after the add/drop period, helping balance the conditions across delivery mode and class size.
The design covered a 16-week semester, with the chatbot intervention operating for 12 calendar weeks. Before treatment, all students completed the same course work and first test. The control group then continued receiving participation credit for study sessions using custom Quizlet flash cards. The treatment group received equivalent credit for at least one chatbot session per chapter, while retaining access to Quizlet without additional credit. Students in both groups could use their assigned support as often as they wished. This made the comparison closer to adding a course chatbot to realistic study options than to replacing instruction.
The chatbot was more carefully constructed than a general-purpose chat window. It ran in a secure university Microsoft Azure environment using Copilot with GPT-4o. The RAG content included instructor-provided concepts, glossaries, learning objectives, and lecture transcripts organized by chapter, but not test questions or textbook content. System instructions gave the tutor goals, a conversational role, and functions for discussion, explanation, multilingual interaction, and short quizzes. The researchers tested its consistency and accuracy and reported that it answered 199 of 200 course test questions correctly, even though those questions were not included in its knowledge base.
The outcomes came from several sources. Students completed pre- and post-treatment measures of individual interest and self-efficacy. Engagement included emotional, participation, performance, and skill scales, weekly self-reports, electronic-book usage, and chatbot session counts. Academic achievement was represented by standardized scores from a common first test before treatment and a common fourth test at the end of the term. The analyses used difference-in-differences models and controlled for gender, age, race, weekly job hours, and whether the course was face-to-face or asynchronous online.
Across interest, self-efficacy, engagement, and test scores, the critical treatment-by-time interactions were not statistically significant. Self-efficacy rose slightly over the semester for students as a whole, and several engagement measures fell, but neither pattern was attributable to chatbot access. For achievement, the treatment-by-time interaction was also non-significant. In other words, the study did not find that adding the course-grounded chatbot changed the measured trajectory relative to the control condition.
Students nevertheless viewed the tool positively. Treatment students reported high enjoyment, perceived help with course material, and interest in having a similar chatbot available in other courses. Their own ratings were more cautious about whether the chatbot improved grades or interest in marketing. This gap is one of the paper's most useful findings: liking an AI tutor, finding it convenient, or wanting continued access does not establish that it improves learning. Adoption and educational effectiveness must be evaluated separately.
The null result also has boundaries. The study involved one introductory subject, one instructor, and three sections at one university. The control group had a legitimate study aid, which sets a more demanding comparison than no support. The chatbot did not remember prior sessions, limiting personalization and social continuity. The researchers could not fully capture how students changed their broader study habits, and the final test covered different course material from the baseline test even though scores were standardized. More advanced learners, other subjects, alternative tutoring strategies, or stronger integration with classroom activity could produce different results.
For AIEDHK, the practical lesson is to evaluate a designed intervention before scaling a platform contract. A course-grounded chatbot may be accurate and popular yet still add no measurable benefit to existing study support. Pilots should define the expected mechanism, compare against a credible alternative, measure independent achievement and engagement over time, inspect actual usage and substitution effects, and test whether memory or personalization helps without creating unacceptable privacy costs. The paper does not show that RAG tutors can never work. It shows that grounding and availability alone are insufficient evidence of learning value.


