
Model Monitoring, Drift, and Incident Response
How educational AI teams detect changing data, behaviour, performance, access, costs, and harms, and how clear thresholds, ownership, containment, communication, remedy, and learning turn monitoring into accountable operations.
Fuentes
Resumen completo
An AI system that passed evaluation before launch can become less suitable over time. Models, prompts, retrieval sources, interfaces, users, courses, and policies change. Monitoring is the planned collection and interpretation of evidence about whether the system continues to work as intended and whether new harms appear. It should begin from the educational purpose and risk assessment, not from whatever metrics a vendor happens to expose.
Drift describes change that can weaken an earlier evaluation. Data drift occurs when inputs change, such as a new curriculum, language mix, or student population. Concept drift occurs when the relationship between inputs and the target changes. Model or product updates can alter output without a local code change. Retrieval sources may become stale. Human practice can drift when users develop shortcuts or use a tool for decisions outside its approved scope. Cost and latency changes can also make access inequitable.
Monitoring combines technical and human signals. Teams can track error rates, calibration, source fidelity, harmful outputs, latency, usage, override rates, accessibility failures, complaints, subgroup patterns, and unaided learning outcomes. Averages need examples and uncertainty. A stable metric does not prove that unmeasured harms are absent. Learners and educators need easy routes to report problems, and reports should be protected from retaliation and reviewed by people able to act.
An incident is an event that threatens safety, rights, learning, security, privacy, or reliable operation. Response plans name severity levels, owners, communication channels, evidence-preservation rules, and authority to pause a feature. Teams contain the issue, protect affected people, investigate causes, correct or roll back, communicate honestly, and provide remedy. Secrets and personal data should not be copied into public logs. A post-incident review examines system and process causes rather than searching only for individual blame.
In education, learners can monitor a fictional writing-feedback model across two terms. They receive charts showing rising citation errors after a model update, slower performance on older devices, and complaints from multilingual students. Groups decide what crosses a stop threshold, what evidence to preserve, whom to notify, and how to support affected learners. They design a rollback and a test for safe return, then publish a plain-language incident summary.
Monitoring is meaningful only when evidence changes decisions. Institutions should define baselines, thresholds, review frequency, update triggers, retention, and escalation before launch. They should periodically retest with representative cases and retire systems whose benefits no longer justify their risks or burden. Drift is not proof of negligence; ignoring observable drift is a governance failure. Responsible operations make change, uncertainty, incidents, and repair visible throughout the life of educational AI. A public monitoring note can summarize the current version, checks, known limitations, recent incidents, corrective actions, owner, and next review without exposing personal or security-sensitive data.


