
An eliciting LLM dashboard prompted more reflective dialogue and showed a trend toward stronger learning-judgment calibration
Laura Graf, Patrick Bassner, Maximilian Anzinger, Felix Dietrich, Stephan Krusche, Oleksandra Poquet
Education and Information Technologies
ملخص 500 كلمة

Graf and colleagues rethink a learning dashboard as an interactive space rather than a static display. Their July 2026 Education and Information Technologies article reports a five-week exploratory case study in an introductory algorithms and data-structures course at a European university. Thirty volunteer computer-science students were compensated and randomly assigned to one of three conditions: a dashboard without a pedagogical agent, an agent that mainly told students information, or an agent designed to elicit reflection through questions. The small sample makes the work exploratory despite the randomized assignment.
The dashboard combined learning analytics with a GPT-4o pedagogical agent and a Judgment of Learning prompt. Students first estimated their own understanding; the interface then revealed a system estimate derived from course activity and performance. The agent could use tools to retrieve exercises, scores, timestamps, competency information and lecture slides. Its ReAct-style process could issue several tool requests for one message. The telling and eliciting conditions therefore differed principally in dialogue strategy, not simply in whether an LLM or analytics data were present.
Reflective messages appeared more often in the eliciting condition than in the telling condition. The reported two-proportion test was z = 2.47, p = .013. This supports a bounded interaction claim: asking learners to explain and inspect their thinking changed the observed dialogue. It does not establish that students acquired more algorithmic knowledge. Static analytics alone were not highly salient to many participants; 70 percent said they did not actively attend to the dashboard visualizations while making their self-rating.
The eliciting condition recorded 104 learning judgments. Those judgments correlated with confidence, progress and system-estimated mastery, with the mastery relationship reaching r = .482, p < .001, in the third interval. The telling condition did not show significant overall relationships on the same measures. Curiosity was also important: 83 percent of participants said wanting to see the system rating motivated them to submit a judgment. These patterns suggest developing calibration, but multiple observations came from a small number of learners and should not be treated as independent proof of improvement.
System quality was imperfect. In an audit of 284 randomly sampled agent responses, 13 percent were rated faulty and 37 percent very useful; during the first two weeks, 37 percent were faulty. The study lasted only five weeks, involved one computer-science course and used a self-selected, paid sample. It did not directly measure changed study behavior, course examination gains, delayed retention or transfer. The findings therefore support more testing of elicitation and calibration, not a claim that an LLM dashboard improves learning outcomes.
For Hong Kong universities and secondary schools, a useful replication could ask learners to explain a mastery judgment before any system score is shown, then test whether that explanation predicts and improves later unaided work. The dashboard should be evaluated in Cantonese, English and Putonghua, with curriculum-grounded retrieval, teacher review and visible correction routes for faulty responses. Engagement, calibration and achievement should be reported separately. The design is promising because it turns analytics into a conversation, but educational value depends on accurate support and evidence beyond the conversation itself.


