用博弈论建模AI安全部署,找更可靠的控制协议。
Games for AI Control: Models of Safety Evaluations of AI Deployment Protocols
- 将安全评估转化为多目标部分可观测博弈,形式化红队对抗过程。
- 可找到帕累托最优的部署协议,比经验研究更优。
- 适合研究AI控制、安全评估的学者与工程师使用。
为评估不可信AI部署协议的安全性与实用性,AI Control采用协议设计者与对手之间的红队对抗。本文提出AI-Control Games,一种形式化的多目标、部分可观测随机博弈模型,用于描述红队对抗过程。我们还建立了AI-Control Games到零和部分可观测随机博弈特例的归约,从而可利用现有算法求解帕累托最优协议。我们将该形式化方法应用于建模、评估与合成部署不可信语言模型作为编程助手的协议,重点聚焦于可信监控协议(Trusted Monitoring),该协议使用较弱的语言模型并结合有限的人类协助。为展示其效用,我们实现了在现有场景中的改进,拓展至新场景的协议评估,并分析了建模假设对协议安全性与实用性的影。最后,我们借助该形式化框架精确揭示了先前控制工作中的一些隐含假设。
原文摘要 · Abstract (English)
To evaluate the safety and usefulness of deployment protocols for untrusted AIs, AI Control uses a red-teaming exercise played between a protocol designer and an adversary. This paper introduces AI-Control Games, a formal decision-making model of the red-teaming exercise as a multi-objective, partially observable, stochastic game. We also introduce reductions from AI-Control Games to a special case of zero-sum partially observable stochastic games that allow us to leverage existing algorithms to find Pareto-optimal protocols. We apply our formalism to model, evaluate and synthesise protocols for deploying untrusted language models as programming assistants, focusing on Trusted Monitoring protocols, which use weaker language models and limited human assistance. To demonstrate the utility of our formalism, we show improvements over empirical studies in existing settings, evaluate protocols in new settings, and analyse how modelling assumptions affect the safety and usefulness of protocols. Finally, we leverage our formalism to precisely describe some of the implicit assumptions in prior control work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。