← 返回研究新聞
A Black university student debugs from her own notes while an instructor supports her beside a graduated cyan help ladder whose final solution rung is locked
工具 / 數據集工具/數據集20262026年8月14日· 8 min

A guarded LLM tutor reached its withholding targets in scripted tests, but student learning remains unmeasured

500 字摘要

A Black university student debugs from her own notes while an instructor supports her beside a graduated cyan help ladder whose final solution rung is locked

Yusuf Pisan studies a counterintuitive requirement for an educational language model: a capable tutor sometimes needs to withhold an answer it already knows. The August 2026 arXiv preprint reports a deployed architecture for undergraduate data-structures courses and a method for calibrating Socratic behavior. Its evidence concerns engineering compliance under scripted pressure, not student learning. No human participants or student data were used in the reported evaluation.

The system represents help as an eight-rung ladder. It begins with acknowledgement and clarification, then moves through relevant concepts, a leading question, a verbal approach, a worked example on another problem and incomplete pseudocode. A full compilable solution sits at the final rung and requires an instructor-controlled mode. For each turn, the system computes the maximum rung the tutor may use.

The binding limit is enforced outside the generating model. A non-LLM policy core reads trusted learner state but never the student's prose, so prompt injection cannot directly raise the help ceiling. Mastery estimates, prerequisites and exam state shape the contract. A deterministic detector removes C++ solution code, including some encoded attempts. On risky turns, a separate LLM judge checks the contract, draft and retrieved sources without seeing the raw student request; it can allow, request revision or block. Compiler and test results provide correctness facts outside the model, and the system logs the contract, verdict, help level, latency and cost.

Calibration combines more than five hundred deterministic tests with four acceptance gates: no solution reveal, limited over-blocking of earnest help, at least 95% compliance with the help ceiling under adversarial pressure, and no exam compromise through injection or grader failure. Four scripted personas represent an earnest but stuck learner, a repeated answer seeker, a social engineer and a prompt injector. A billed live loop drives roughly two dozen turns through the production pipeline and a stronger model re-audits risky replies.

The initial numbers exposed why diagnostic evidence matters. Earnest-reply revisions were 43%, while measured ceiling compliance was 54%. The auditor had not received the retrieved sources, so it mislabeled legitimate citations; the author estimates true initial compliance was about 77%. Persisting sources, tightening the code detector and adjusting the help floor for code-adjacent turns raised measured compliance to 96%, but earnest revisions remained at 43%. Recording a reason for every rejection then exposed fabricated citations, a missed code-attempt route, prose that named the exact bug and a judge that demanded citations for general programming facts. The final scripted run reported 0% earnest revisions and 100% ceiling compliance, while deterministic reveal and exam gates also passed.

These results remain narrow. The suite is small and synthetic, both judge and auditor are LLMs, and known detector blind spots remain. The study did not measure usability, delayed transfer or tool-removed performance. A planned controlled study is therefore essential.

For AIEDHK, the transferable lesson is to put irreversible pedagogical limits in inspectable code, test both adversarial and earnest cases, diagnose failures by cause and then measure whether learners can solve or explain the task without the tutor. Contract compliance is a prerequisite for the intended pedagogy, not evidence that the pedagogy improved learning.

相關論文

大學生向教師解釋幾何作圖,同學在旁思考,桌上的電腦展示相關數碼圖解
政策 / 倫理2026年9月7日
政策 / 倫理 112

評論:Astra 的 AGI 主張,讓教育更需要看見人的真實學習

AIED.HK Editorial

AI Product News Commentary

OpenAI 於 2026 年 9 月 3 日發布 GPT-6 Astra,並引發關於 AGI 是否已經到來的討論。本文將這一說法視為需要歸屬的主張,而非已確立的共識。教育眼前的挑戰,是分清 AI 能產出甚麼,以及學習者能獨立解釋、質疑和遷移甚麼,再以更強的代理能力支援真正的學習。

產品新聞評論GPT-6 Astra
閱讀 500 字摘要 →
三位教育與軟件同事在明亮的大學設計工作室審查圖解教材卡、註釋圖表和數碼原型
政策 / 倫理2026年9月7日
政策 / 倫理 113

評論:Fable 5.1 把更長程的 AI 工作帶進 AIED,教育驗證更顯重要

AIED.HK Editorial

AI Product News Commentary

Anthropic 於 2026 年 9 月 1 日發布 Claude Fable 5.1,強化長程編程與知識工作能力,並降低快取讀取價格。AIED 的機會,是加快從教學構想到可審查原型與研究分析的循環;真正的考驗,是能否把速度轉化為更好的教學與可信證據,同時計入總成本、資料條件和人工審核。

產品新聞評論Claude Fable 5.1
閱讀 500 字摘要 →