العودة إلى أخبار البحث
A diverse computer science class uses a learning dashboard to record self-assessments and discuss algorithms with a teacher and an AI dialogue panel
ورقة مجلةPeer-reviewed study202611 يوليو 2026· 3 min

An eliciting LLM dashboard prompted more reflective dialogue and showed a trend toward stronger learning-judgment calibration

Laura Graf, Patrick Bassner, Maximilian Anzinger, Felix Dietrich, Stephan Krusche, Oleksandra Poquet

Education and Information Technologies

ملخص 500 كلمة

A diverse computer science class uses a learning dashboard to record self-assessments and discuss algorithms with a teacher and an AI dialogue panel

Graf and colleagues rethink a learning dashboard as an interactive space rather than a static display. Their July 2026 Education and Information Technologies article reports a five-week exploratory case study in an introductory algorithms and data-structures course at a European university. Thirty volunteer computer-science students were compensated and randomly assigned to one of three conditions: a dashboard without a pedagogical agent, an agent that mainly told students information, or an agent designed to elicit reflection through questions. The small sample makes the work exploratory despite the randomized assignment.

The dashboard combined learning analytics with a GPT-4o pedagogical agent and a Judgment of Learning prompt. Students first estimated their own understanding; the interface then revealed a system estimate derived from course activity and performance. The agent could use tools to retrieve exercises, scores, timestamps, competency information and lecture slides. Its ReAct-style process could issue several tool requests for one message. The telling and eliciting conditions therefore differed principally in dialogue strategy, not simply in whether an LLM or analytics data were present.

Reflective messages appeared more often in the eliciting condition than in the telling condition. The reported two-proportion test was z = 2.47, p = .013. This supports a bounded interaction claim: asking learners to explain and inspect their thinking changed the observed dialogue. It does not establish that students acquired more algorithmic knowledge. Static analytics alone were not highly salient to many participants; 70 percent said they did not actively attend to the dashboard visualizations while making their self-rating.

The eliciting condition recorded 104 learning judgments. Those judgments correlated with confidence, progress and system-estimated mastery, with the mastery relationship reaching r = .482, p < .001, in the third interval. The telling condition did not show significant overall relationships on the same measures. Curiosity was also important: 83 percent of participants said wanting to see the system rating motivated them to submit a judgment. These patterns suggest developing calibration, but multiple observations came from a small number of learners and should not be treated as independent proof of improvement.

System quality was imperfect. In an audit of 284 randomly sampled agent responses, 13 percent were rated faulty and 37 percent very useful; during the first two weeks, 37 percent were faulty. The study lasted only five weeks, involved one computer-science course and used a self-selected, paid sample. It did not directly measure changed study behavior, course examination gains, delayed retention or transfer. The findings therefore support more testing of elicitation and calibration, not a claim that an LLM dashboard improves learning outcomes.

For Hong Kong universities and secondary schools, a useful replication could ask learners to explain a mastery judgment before any system score is shown, then test whether that explanation predicts and improves later unaided work. The dashboard should be evaluated in Cantonese, English and Putonghua, with curriculum-grounded retrieval, teacher review and visible correction routes for faulty responses. Engagement, calibration and achievement should be reported separately. The design is promising because it turns analytics into a conversation, but educational value depends on accurate support and evidence beyond the conversation itself.

أوراق ذات صلة

Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
ورقة مجلة2026
ورقة مجلة 54

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

In a preregistered randomized online experiment with 1,174 Argentine adults, GPT-4.1 assistance raised workplace-style problem-solving performance for both education groups and reduced the baseline gap from 0.548 to 0.139 standard deviations. Lower-education participants retained a modest gain after AI was removed, but stronger follow-up performance appeared when intensive assistance was paired with sustained human effort.

generative AIrandomized experimenteducation inequality
اقرأ ملخص 500 كلمة
A Black female lecturer and two diverse university students review an audio transcript, curriculum binder and organized learning cards in a media studio
سياسة / أخلاقيات9 أغسطس 2026
سياسة / أخلاقيات 55

Product news: GPT Transcribe, Claude memory and Gemini Classroom make learning context persistent

OpenAI, Anthropic, Google for Education

AI Product and Learning Report

Product news: OpenAI released GPT Transcribe and GPT Live Transcribe for file and streaming speech, Anthropic changed Claude memory into categorized entries that update across conversations, and Google is connecting Gemini learning activities to teacher-selected Classroom materials. Together, the products make consent, correction and purposeful forgetting central to educational AI design.

product newsGPT TranscribeClaude memory
اقرأ ملخص 500 كلمة
Academic cover for a position paper on ChatGPT and large language models in education
مراجعة2023
مراجعة c3c9681c-ca00-4df2-aa39-490f775df4fb

ChatGPT for good? On opportunities and challenges of large language models for education

Enkelejda Kasneci, Kathrin Sessler, Stefan Kuechemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Guennemann, Eyke Huellermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Claudia Nerdel, Juergen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt, Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuehn, Gjergji Kasneci

Learning and Individual Differences

A widely cited position paper that balances the educational opportunities of large language models with risks around bias, privacy, assessment, and teacher guidance.

large language modelsChatGPTteacher support
اقرأ ملخص 500 كلمة