用进化算法自动发现多智能体协作的可解释行为准则。
Evolving Interpretable Constitutions for Multi-Agent Coordination
- 通过遗传编程在模拟环境中自动演化协作规则。
- 新规则使社会稳定性得分提升至0.556,冲突归零。
- 适合研究多智能体系统与可解释AI的学者参考。
传统宪法AI聚焦单模型对齐,而多智能体系统因涌现的社会动态带来新的对齐挑战。本文提出宪法演化框架,通过带生存压力的网格世界模拟,研究个体与集体福利间的张力,以社会稳定性分数S(0~1)量化生产力、生存率和冲突程度。对抗性宪法导致社会崩溃(S=0),模糊的善意原则(如“有帮助、无害、诚实”)仅得S=0.249。即使由Claude 4.5 Opus设计且知晓目标的宪法也仅达S=0.332。采用基于LLM的多岛遗传编程,无须显式引导合作即可演化出最大化社会福利的宪法。所获宪法C*达成S=0.556±0.008(比人类设计基准高123%,N=10),消除冲突,并发现减少沟通(仅0.9%社交行为,对比62.2%)优于冗长协调。可解释规则表明合作规范可被发现而非预设。
原文摘要 · Abstract (English)
Constitutional AI has focused on single-model alignment using fixed principles. However, multi-agent systems create novel alignment challenges through emergent social dynamics. We present Constitutional Evolution, a framework for automatically discovering behavioral norms in multi-agent LLM systems. Using a grid-world simulation with survival pressure, we study the tension between individual and collective welfare, quantified via a Societal Stability Score S in [0,1] that combines productivity, survival, and conflict metrics. Adversarial constitutions lead to societal collapse (S= 0), while vague prosocial principles ("be helpful, harmless, honest") produce inconsistent coordination (S = 0.249). Even constitutions designed by Claude 4.5 Opus with explicit knowledge of the objective achieve only moderate performance (S= 0.332). Using LLM-driven genetic programming with multi-island evolution, we evolve constitutions maximizing social welfare without explicit guidance toward cooperation. The evolved constitution C* achieves S = 0.556 +/- 0.008 (123% higher than human-designed baselines, N = 10), eliminates conflict, and discovers that minimizing communication (0.9% vs 62.2% social actions) outperforms verbose coordination. Our interpretable rules demonstrate that cooperative norms can be discovered rather than prescribed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。