Назад к новостям исследований
A learning scientist and two diverse student researchers compare coded learner actions with a structured reasoning map in a university lab
Конференционная статьяConference paper202611 авг. 2026 г.· 2 min

INSIDE aligned simulated student actions with internal dialogue, but reasoning fidelity remained partial

Rose Niousha, Minwoo Kang, Narges Norouzi

Conference on Language Modeling 2026

Резюме на 500 слов

A learning scientist and two diverse student researchers compare coded learner actions with a structured reasoning map in a university lab

Rose Niousha, Minwoo Kang and Narges Norouzi examine a hidden weakness in large-language-model student simulators. A simulator may reproduce an observable answer or code submission while representing the learner's reason incorrectly. Two students can make the same mistake because of different knowledge, goals, emotions, or strategies. A tutoring system tested only against simulated actions could therefore appear effective for the wrong reason.

The authors introduce INTERNAL STUDENT DIALOGUE, or INSIDE, a framework that fine-tunes models to generate both latent dialogue and student action. The dialogue is structured through Bloom's taxonomy across cognitive, affective, and action dimensions. Training examples pair think traces with observable outputs. Evaluation then considers two axes: whether simulated actions resemble those of real students and whether the generated internal dialogue aligns with available evidence about learner reasoning.

Across the evaluated models, INSIDE improved action fidelity and produced the strongest reported reasoning alignment, reaching up to 57.9 percent. This matters because it moves evaluation beyond surface imitation. A simulator that states a plausible misconception or uncertainty can help researchers inspect why a tutor chooses a hint and design scenarios that vary more meaningfully than correct versus incorrect answers. Joint modelling can also reveal when a convincing action is paired with an implausible internal explanation.

The result should not be read as access to a student's mind. Internal dialogue is a constructed representation, and reasoning alignment below full agreement leaves substantial mismatch. Bloom categories organize aspects of performance but do not capture every cultural, social, emotional, or contextual influence on learning. Fine-tuned models may reproduce biases in the traces used for training. A simulator can also generate a coherent rationale that no real student expressed. Acceptance at COLM supports research relevance, not classroom validity.

For tutor evaluation, simulated learners are best used as a pre-deployment stress test. Researchers can generate cases with different misconceptions, confidence, affect, and action patterns, then inspect whether the tutor adapts safely. Results should guide hypotheses and identify failure modes before real-user study. They should not replace participatory design, teacher review, or trials with actual learners. Sensitive decisions about intervention, grading, or disability support should never be based on a synthetic profile treated as a real person's state.

For AIEDHK, INSIDE offers a valuable methodological warning and a promising tool. Observable fidelity alone is insufficient when an educational system claims to respond to reasoning. Researchers should report how latent states were defined, whose data informed them, how alignment was measured, and where the simulator fails. They should then validate tutor behavior with diverse learners and outcome evidence. A simulator can make early testing broader and safer, but the remaining gap between generated dialogue and human experience must remain visible. The strongest use is to challenge a tutor before deployment, not to certify that it understands students. Evaluation sets should include multilingual, neurodiverse, and culturally varied reasoning patterns without presenting any synthetic profile as a definitive account of a group.

Связанные статьи

A trained undergraduate tutor supports two learners from different disciplines as they frame an AI project and verify a working prototype in a university studio
Конференционная статья2026
Конференционная статья 108

LearnAI piloted a two-layer route from campus AI awareness to supervised co-creation

Weihao Qu, Ling Zheng, Chris Buzaid, Daniel Crawford

arXiv preprint

Qu and colleagues describe LearnAI, a university framework that paired short course-embedded introductions across 18 courses with optional one-to-one co-creation led by trained undergraduate tutors. Thirty-five clients produced 36 portfolios and more than 20 deployed applications. Interviews and a seven-person readiness comparison are preliminary, so the report supports feasibility and design learning rather than causal claims about skill growth.

AI co-creationpeer tutoringcross-disciplinary learning
Читать резюме на 500 слов
A diverse university project team reviews a compact result panel, a gateway diagram, and a verification checklist in a bright computing studio
Политика / этика21 авг. 2026 г.
Политика / этика 109

Product news: Claude Code 2.1.237 adds a concise output style and repairs gateway prompt caching

Anthropic

AI Product and Learning Report

Product news: Claude Code 2.1.237 introduces a built-in Concise output style and fixes prompt caching for sessions that use an LLM gateway or custom base URL. The release can reduce narration and repeated processing, but brevity and cache efficiency do not establish correctness. Educational teams should preserve task requirements, evidence, tests, and review notes outside the presentation style.

product newsClaude Code 2.1.237output styles
Читать резюме на 500 слов
A diverse group of computing students and an instructor compare AI tutoring responses with a prerequisite map and scaffolded Python exercises
Конференционная статья2026
Конференционная статья 86

Pedagogical-fit feedback improved 51 of 62 weak AI-tutor responses, but scaffolding gains created trade-offs

Benjamin Barlog, Hudson Craig, Zedong Peng

IEEE 27th International Conference on Information Reuse and Integration for Data Science (IRI 2026)

Barlog, Craig and Peng evaluated ChatGPT, Gemini, Gemma 4 and Qwen 3 across 240 introductory-programming tutoring scenarios with a six-part Pedagogical Suitability Index. Baseline scores differed modestly, while targeted feedback improved 51 of 62 weak cases. The largest gain was scaffolding, but prerequisite ordering and Bloom-level alignment sometimes declined, showing why tutoring quality needs multiple measures rather than one composite score.

AI tutorspedagogical fitscaffolding
Читать резюме на 500 слов