
INSIDE aligned simulated student actions with internal dialogue, but reasoning fidelity remained partial
Rose Niousha, Minwoo Kang, Narges Norouzi
Conference on Language Modeling 2026
Résumé de 500 mots

Rose Niousha, Minwoo Kang and Narges Norouzi examine a hidden weakness in large-language-model student simulators. A simulator may reproduce an observable answer or code submission while representing the learner's reason incorrectly. Two students can make the same mistake because of different knowledge, goals, emotions, or strategies. A tutoring system tested only against simulated actions could therefore appear effective for the wrong reason.
The authors introduce INTERNAL STUDENT DIALOGUE, or INSIDE, a framework that fine-tunes models to generate both latent dialogue and student action. The dialogue is structured through Bloom's taxonomy across cognitive, affective, and action dimensions. Training examples pair think traces with observable outputs. Evaluation then considers two axes: whether simulated actions resemble those of real students and whether the generated internal dialogue aligns with available evidence about learner reasoning.
Across the evaluated models, INSIDE improved action fidelity and produced the strongest reported reasoning alignment, reaching up to 57.9 percent. This matters because it moves evaluation beyond surface imitation. A simulator that states a plausible misconception or uncertainty can help researchers inspect why a tutor chooses a hint and design scenarios that vary more meaningfully than correct versus incorrect answers. Joint modelling can also reveal when a convincing action is paired with an implausible internal explanation.
The result should not be read as access to a student's mind. Internal dialogue is a constructed representation, and reasoning alignment below full agreement leaves substantial mismatch. Bloom categories organize aspects of performance but do not capture every cultural, social, emotional, or contextual influence on learning. Fine-tuned models may reproduce biases in the traces used for training. A simulator can also generate a coherent rationale that no real student expressed. Acceptance at COLM supports research relevance, not classroom validity.
For tutor evaluation, simulated learners are best used as a pre-deployment stress test. Researchers can generate cases with different misconceptions, confidence, affect, and action patterns, then inspect whether the tutor adapts safely. Results should guide hypotheses and identify failure modes before real-user study. They should not replace participatory design, teacher review, or trials with actual learners. Sensitive decisions about intervention, grading, or disability support should never be based on a synthetic profile treated as a real person's state.
For AIEDHK, INSIDE offers a valuable methodological warning and a promising tool. Observable fidelity alone is insufficient when an educational system claims to respond to reasoning. Researchers should report how latent states were defined, whose data informed them, how alignment was measured, and where the simulator fails. They should then validate tutor behavior with diverse learners and outcome evidence. A simulator can make early testing broader and safer, but the remaining gap between generated dialogue and human experience must remain visible. The strongest use is to challenge a tutor before deployment, not to certify that it understands students. Evaluation sets should include multilingual, neurodiverse, and culturally varied reasoning patterns without presenting any synthetic profile as a definitive account of a group.


