多个不完全对齐的AI竞争,也能让人类获得接近完美对齐的效果。
Emergent Alignment via Competition
- 通过多智能体博弈机制,利用竞争促使弱对齐模型协同逼近最优决策。
- 在凸包条件下,用户在所有均衡中都能实现接近贝叶斯最优的决策效果。
- 即使用户只选表现最好的单一模型,也能保持近优性能,无需额外假设。
将人工智能系统与人类价值观对齐仍是根本性挑战,但无法构建完全对齐模型是否意味着无法获得对齐带来的好处?我们研究一种情境:人类用户与多个不同程度偏移的AI代理交互,这些代理均非完全对齐。关键洞察是,当用户的效用位于各代理效用的凸包内时——这一条件随模型多样性增加而更易满足——战略竞争可产生与与完美对齐模型交互相当的结果。我们将此建模为多领导者斯塔克尔伯格博弈,将贝叶斯劝说扩展至多方信息不对称的多轮对话,并证明三个结论:(1) 当完美对齐允许用户学习其贝叶斯最优行动时,在凸包条件下用户在所有均衡中均可实现该目标;(2) 在较弱假设下仅需近似效用学习,采用量化响应的非策略用户在所有均衡中可获得近最优效用;(3) 当用户在评估期后选择表现最佳的单一模型时,均衡保证仍为近最优,且无需进一步分布假设。理论之外,我们通过两组实验进行补充验证。
原文摘要 · Abstract (English)
Aligning AI systems with human values remains a fundamental challenge, but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? We study a strategic setting where a human user interacts with multiple differently misaligned AI agents, none of which are individually well-aligned. Our key insight is that when the users utility lies approximately within the convex hull of the agents utilities, a condition that becomes easier to satisfy as model diversity increases, strategic competition can yield outcomes comparable to interacting with a perfectly aligned model. We model this as a multi-leader Stackelberg game, extending Bayesian persuasion to multi-round conversations between differently informed parties, and prove three results: (1) when perfect alignment would allow the user to learn her Bayes-optimal action, she can also do so in all equilibria under the convex hull condition (2) under weaker assumptions requiring only approximate utility learning, a non-strategic user employing quantal response achieves near-optimal utility in all equilibria and (3) when the user selects the best single AI after an evaluation period, equilibrium guarantees remain near-optimal without further distributional assumptions. We complement the theory with two sets of experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。