arXiv:2507.03525cs.AIcs.SY2025-07被引 3

厘清AI监管中监督与控制的区别,提出可落地的评估框架。

Limits of Safe AI Deployment: Differentiating Oversight and Control

  • 区分事前/实时控制与事后监督,明确功能差异
  • 提出符合技术与组织现实的监管对齐框架
  • 给出监督机制适用边界,助力监管与实践决策

监督(包含控制与督导)常被视作确保AI系统可问责、可靠并满足治理要求的关键。然而,‘人类监督’等术语在法规中使用时存在模糊性,易导致概念理解不一,削弱系统设计与评估的有效性。本文针对非AI领域的监督文献进行批判性回顾,并简要梳理了相关AI研究。进一步区分控制为事前或实时操作性机制,而监督则为事后政策与治理职能;控制旨在预防失败,监督则聚焦于检测、补救及未来预防激励。据此提出三项贡献:1)构建框架,将监管期望与技术及组织可行性对齐,明确各机制的实现条件、局限及实际意义所需要素;2)建议监督方法应纳入风险管理文档,并基于微软负责任AI成熟度模型,提出AI监督成熟度模型;3)明确界定这些机制的应用边界、失效场景以及现有方法无法覆盖的领域。该工作有助于判断特定部署情境下是否存在有意义的监督,并支持监管者、审计者与从业者识别当前与未来的局限。

原文摘要 · Abstract (English)

Oversight and control, which we collectively call supervision, are often discussed as ways to ensure that AI systems are accountable, reliable, and able to fulfill governance and management requirements. However, the requirements for "human oversight" risk codifying vague or inconsistent interpretations of key concepts like oversight and control. This ambiguous terminology could undermine efforts to design or evaluate systems that must operate under meaningful human supervision. This matters because the term is used by regulatory texts such as the EU AI Act. This paper undertakes a targeted critical review of literature on supervision outside of AI, along with a brief summary of past work on the topic related to AI. We next differentiate control as ex-ante or real-time and operational rather than policy or governance, and oversight as performed ex-post, or a policy and governance function. Control aims to prevent failures, while oversight focuses on detection, remediation, or incentives for future prevention. Building on this, we make three contributions. 1) We propose a framework to align regulatory expectations with what is technically and organizationally plausible, articulating the conditions under which each mechanism is possible, where they fall short, and what is required to make them meaningful in practice. 2) We outline how supervision methods should be documented and integrated into risk management, and drawing on the Microsoft Responsible AI Maturity Model, we outline a maturity model for AI supervision. 3) We explicitly highlight boundaries of these mechanisms, including where they apply, where they fail, and where it is clear that no existing methods suffice. This foregrounds the question of whether meaningful supervision is possible in a given deployment context, and can support regulators, auditors, and practitioners in identifying both present and future limitations.

AI监管监督机制风险治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。