提出多智能体系统中可纠正性的博弈框架,解决智能体如何主动求助人类的问题。
On Corrigibility and Alignment in Multi Agent Games
- 将可纠正性建模为双人贝叶斯博弈,让每个智能体都有向人类求助的行动选项。
- 在对称与对抗场景中证明:当智能体对人类偏好和理性程度有合理信念时,可诱导出可纠正行为。
- 适用于需要安全交互的多智能体系统,如协作机器人、自动化决策平台。
自主智能体的可纠正性是系统设计中未被充分探索的部分,以往研究主要聚焦于单智能体系统。已有研究表明,对人类偏好的不确定性有助于保持智能体的可纠正性,即使面对人类非理性行为。本文提出一个通用框架,将多智能体环境中的可纠正性建模为两人博弈,其中双方智能体始终拥有向人类请求监督的动作。该框架采用贝叶斯博弈形式,引入对人类信念的不确定性。进一步分析两种具体情形:第一种是双智能体可纠正性博弈,要求在共同收益(单调)游戏和调和游戏下,两个智能体均表现出可纠正性;第二种是对手设置,其中一个智能体为“防御”方,另一个为“对手”。文中给出了一个通用结论:防御方必须具备关于博弈结构和人类理性的何种信念,才能诱导出可纠正行为。
原文摘要 · Abstract (English)
Corrigibility of autonomous agents is an under explored part of system design, with previous work focusing on single agent systems. It has been suggested that uncertainty over the human preferences acts to keep the agents corrigible, even in the face of human irrationality. We present a general framework for modelling corrigibility in a multi-agent setting as a 2 player game in which the agents always have a move in which they can ask the human for supervision. This is formulated as a Bayesian game for the purpose of introducing uncertainty over the human beliefs. We further analyse two specific cases. First, a two player corrigibility game, in which we want corrigibility displayed in both agents for both common payoff (monotone) games and harmonic games. Then we investigate an adversary setting, in which one agent is considered to be a `defending' agent and the other an `adversary'. A general result is provided for what belief over the games and human rationality the defending agent is required to have to induce corrigibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。