提出智能违抗博弈框架,解决助手何时该违抗指令以保障安全。
The Intelligent Disobedience Game: Formulating Disobedience in Stackelberg Games and Markov Decision Processes
- 基于斯塔克尔伯格博弈建模人机交互,考虑信息不对称下的决策策略。
- 发现'安全陷阱'现象:系统长期避险却无法达成人类目标。
- 可转为多智能体马尔可夫决策过程,适合训练安全违抗的强化学习代理。
在共享自主场景中,自动化助手面临关键矛盾:是遵从人类指令,还是故意违背以防止伤害。这种安全关键行为称为智能违抗。本文提出智能违抗博弈(IDG),一种基于斯塔克尔伯格博弈的序贯博弈框架,建模人类领导者与辅助跟随者在信息不对称下的互动。该框架刻画了多步场景中双方的最优策略,识别出如“安全陷阱”等战略现象——系统持续规避伤害但无法实现人类目标。IDG为设计能学习安全非服从性的智能体提供了数学基础,并支持对人类如何感知和信任违抗型AI的实证研究。论文还将IDG转化为共享控制的多智能体马尔可夫决策过程(MADP),形成紧凑的计算测试平台,用于训练强化学习代理。
原文摘要 · Abstract (English)
In shared autonomy, a critical tension arises when an automated assistant must choose between obeying a human's instruction and deliberately overriding it to prevent harm. This safety-critical behavior is known as intelligent disobedience. To formalize this dynamic, this paper introduces the Intelligent Disobedience Game (IDG), a sequential game-theoretic framework based on Stackelberg games that models the interaction between a human leader and an assistive follower operating under asymmetric information. It characterizes optimal strategies for both agents across multi-step scenarios, identifying strategic phenomena such as ``safety traps,'' where the system indefinitely avoids harm but fails to achieve the human's goal. The IDG provides a needed mathematical foundation that enables both the algorithmic development of agents that can learn safe non-compliance and the empirical study of how humans perceive and trust disobedient AI. The paper further translates the IDG into a shared control Multi-Agent Markov Decision Process representation, forming a compact computational testbed for training reinforcement learning agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。