仅靠机制设计无法实现最佳社会福祉,需让AI具备利他特质
Mechanism Design Is Not Enough: Prosocial Agents for Cooperative AI

- 基于不完全契约理论,证明机制无法消除所有福利损失
- 实验显示利他型大模型在资源分配与社会困境中表现更优
- 适合关注大规模协作AI安全的开发者与研究者
确保人工智能代理在与其他主体交互时表现出安全且有益的行为,已成为现代人工智能安全的核心挑战。尽管机制设计(即设计规则以对齐个体与集体目标)可激励合作行为,但其是否足以最大化大型语言模型代理的社会福祉仍存疑问。本文从不完全契约理论出发,正式证明:当契约无法区分所有相关未来情境时,必然存在无法被任何现实机制消除的正向福利损失。我们进一步表明,具备利他特质的代理——即同时考虑自身与他人福祉的代理——能够弥合这一差距,实现社会最优且对个体也更有利的结果。实验结果表明,在多智能体资源分配环境和经典社会困境场景中,由大语言模型驱动的利他型代理表现更佳。对人工智能安全的启示是明确的:要实现大规模协作互动,仅设计良好机制是不够的,必须让代理本身具备内在利他性。
原文摘要 · Abstract (English)
Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety. While mechanism design, as the theory of designing rules to align individual and collective objectives, can incentivize cooperative behavior, it is still an open question whether it alone is sufficient to maximize LLM agents' social welfare. This work proves that the answer is negative: drawing from incomplete contract theory, we formally show that when contracts cannot distinguish all relevant future contingencies, there is a strictly positive welfare loss that no realistic mechanism can eliminate. We show that prosocial agents, who weigh others' welfare alongside their own, can close this gap and achieve outcomes that are socially superior and individually beneficial. Experimentally, we show that in multi-agent resource-allocation environments and canonical social dilemmas where agents are powered by large language models, prosociality is beneficial. The implication for AI safety is clear: to enable cooperative interactions at scale, designing adequate mechanisms is not sufficient; agents must be built to be intrinsically prosocial.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。