arXiv:2412.00033cs.AIcs.CY2024-12NeurIPS被引 1

提出可验证安全的智能体政策,让AI治理政府更可信

Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies

  • 用社会选择理论定义智能体对齐度,量化其与社会目标的接近程度
  • 证明近似对齐策略存在性,并给出其成立的充分条件
  • 设计无需黑箱内部结构的安全保障机制,适合高风险决策场景

尽管自主智能体在处理海量复杂数据方面常超越人类,但其潜在的对齐问题(即真实目标不透明)至今阻碍其在社会决策等关键应用中的使用。现有对齐方法无法为模型安全性提供形式化保证。本文基于效用与社会选择理论,提出社会决策背景下对齐的定量新定义。在此基础上,引入“可能近似对齐”(probabilistically approximately aligned)策略,并推导其存在的充分条件。针对该条件难以满足的实际困难,进一步提出“安全”(safe,即非破坏性)策略概念,设计一种简单而稳健的方法,可对任意黑箱智能体的策略进行保护,确保其所有行为对社会可验证地安全。

原文摘要 · Abstract (English)

While autonomous agents often surpass humans in their ability to handle vast and complex data, their potential misalignment (i.e., lack of transparency regarding their true objective) has thus far hindered their use in critical applications such as social decision processes. More importantly, existing alignment methods provide no formal guarantees on the safety of such models. Drawing from utility and social choice theory, we provide a novel quantitative definition of alignment in the context of social decision-making. Building on this definition, we introduce probably approximately aligned (i.e., near-optimal) policies, and we derive a sufficient condition for their existence. Lastly, recognizing the practical difficulty of satisfying this condition, we introduce the relaxed concept of safe (i.e., nondestructive) policies, and we propose a simple yet robust method to safeguard the black-box policy of any autonomous agent, ensuring all its actions are verifiably safe for the society.

AI对齐社会决策安全策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。