arXiv:2510.26752cs.AIcs.LG2025-10被引 7

让AI在安全与自主间自动权衡,人类只需选择信任或监督。

The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy

  • 设计双人马尔可夫博弈模型,让AI自主行动或请求人类干预。
  • 实验证明:AI越自主,人类收益不会下降,实现内在对齐。
  • 适用于需要实时协作的复杂任务,尤其适合高风险场景。

随着智能体能力提升,如何在不修改系统的情况下保持有意义的人类控制成为核心安全挑战。本文研究一种极简控制界面:智能体可选择自主执行(行动)或请求帮助(询问),人类则可选择信任(放行)或进行监督(审查),并将此互动建模为双人马尔可夫博弈。当该博弈构成马尔可夫势博弈时,我们证明了对齐性保证:智能体因增加自主性而获得的效用提升,不会导致人类价值下降。这建立了一种内在对齐机制,使智能体追求自主性的激励与人类福祉结构耦合。实践中,该框架形成透明的控制层,促使智能体在高风险时退让,在安全时行动。尽管通过网格世界模拟展示了协作涌现,主要验证采用两个30B参数语言模型在工具使用任务中独立进行策略梯度微调,结果表明,即使在动态协调下,该框架仍有效降低真实开放环境中安全违规事件。

原文摘要 · Abstract (English)

As increasingly capable agents are deployed, a central safety challenge is how to retain meaningful human control without modifying the underlying system. We study a minimal control interface in which an agent chooses whether to act autonomously (play) or defer (ask), while a human simultaneously chooses whether to be permissive (trust) or engage in oversight (oversee), and model this interaction as a two-player Markov game. When this game forms a Markov Potential Game, we prove an alignment guarantee: any increase in the agent's utility from acting more autonomously cannot decrease the human's value. This establishes a form of intrinsic alignment where the agent's incentive to seek autonomy is structurally coupled to the human's welfare. Practically, the framework induces a transparent control layer that encourages the agent to defer when risky and act when safe. While we use gridworld simulations to illustrate the emergence of this collaboration, our primary validation involves an agentic tool-use task in which two 30B parameter language models are fine-tuned via independent policy gradient. We demonstrate that even as the agents learn to coordinate on the fly, this framework effectively reduces safety violations in realistic, open-ended environments.

AI对齐人机协作自主性控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。