返回学院
A Black woman educator points to an illustrated risk matrix while an East Asian man holds a red tile and a light-skinned woman holds an amber tile in a bright workshop
AI 知识核心2026年8月12日· 2 分钟

AI Safety and Risk Management

How to identify, evaluate, reduce, monitor, and communicate AI risks across a system's lifecycle while preserving human agency and accountability.

AI safetyrisk managementresponsible AI

来源

完整课程摘要

AI safety is the practice of reducing unacceptable harm from AI systems while enabling worthwhile educational uses in practice. It is not a single technical feature and it does not promise zero risk. A hazard is a source of possible harm, such as disclosure of learner data or a biased recommendation. Risk considers both the likelihood of an event and the severity of its consequences in a particular context. The same model can therefore present different risks when it drafts a low-stakes example, recommends an intervention, or influences access to an educational opportunity.

Risk management begins before a tool is selected. The NIST AI Risk Management Framework organizes continuous work through four connected functions: Govern, Map, Measure, and Manage. Governance establishes accountable roles, policies, documentation, and risk tolerances. Mapping describes the intended purpose, affected people, operating conditions, dependencies, and foreseeable misuse. Measurement tests performance, uncertainty, privacy, security, bias, accessibility, and human-AI interaction with evidence suited to the real setting. Management prioritizes risks, applies controls, decides whether deployment should proceed, and monitors what happens after release.

Educational risks extend beyond an incorrect answer. They can include fabricated sources, unequal error rates, harmful stereotypes, exposure of personal information, insecure integrations, inappropriate content, loss of learner agency, automation bias, inaccessible design, and assessment results that no longer support their intended interpretation. Generative and agentic systems add risks because outputs vary and tools can act on files or services. An agentic loop that gathers context, takes action, and verifies results can be useful, but its permissions, data access, stopping points, and verification evidence must match the stakes.

Effective controls are layered. A school can limit a system to a clearly defined use, minimize the data it receives, test representative scenarios, require human approval for consequential actions, restrict tool permissions, keep useful logs, make actions reversible, and prepare incident and appeal routes. Red-team exercises may reveal failure modes, but one benchmark or safety score cannot represent every learner, language, subject, or future condition. Residual risk should be documented rather than hidden, and controls should be re-evaluated when the model, users, data, or purpose changes.

In education, learners can examine a fictional AI study adviser before it is adopted. They identify stakeholders, intended benefits, possible harms, and evidence gaps; rank each risk by likelihood and consequence; propose controls; and name the person accountable for monitoring and intervention. They then test cases involving ambiguous language, disability access, and sensitive data, record remaining uncertainty, and decide whether to proceed, redesign, limit, or reject the use. This exercise treats safety as disciplined inquiry and shared responsibility. The aim is not fear or blind trust, but proportionate judgment that keeps educational purpose, rights, equity, evidence, and human agency at the center.