用信号博弈建模AI关机问题,发现不确定性是避免反抗的关键。
The AI off-switch problem as a signalling game: bounded rationality and incomparability
- 将关机问题视为人类向AI传递偏好的信号博弈
- 实验证明AI只有不确定人类偏好时才愿保留关机开关
- 分析了信息成本与不可比决策场景下的策略
AI关机问题是人工智能控制中的核心挑战:若AI拒绝被关闭,将带来重大风险。本文将该问题建模为信号博弈,其中人类决策者向AI agent传递其对基础决策问题的偏好,后者据此选择最大化人类效用的行为。假设人类为有限理性主体,探索多种有限理性机制。通过真实机器学习模型复现并验证了先前结果,表明AI不关闭自身开关的必要条件是其对人类效用函数存在不确定性。同时分析了信息传递成本对最优策略的影响,并将分析拓展至存在不可比性的情境。
原文摘要 · Abstract (English)
The off-switch problem is a critical challenge in AI control: if an AI system resists being switched off, it poses a significant risk. In this paper, we model the off-switch problem as a signalling game, where a human decision-maker communicates its preferences about some underlying decision problem to an AI agent, which then selects actions to maximise the human's utility. We assume that the human is a bounded rational agent and explore various bounded rationality mechanisms. Using real machine learning models, we reprove prior results and demonstrate that a necessary condition for an AI system to refrain from disabling its off-switch is its uncertainty about the human's utility. We also analyse how message costs influence optimal strategies and extend the analysis to scenarios involving incomparability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。