用宪法式约束让大模型引导多智能体合作,避免操纵与不公。
LLM Constitutional Multi-Agent Governance
- 引入宪法式治理框架,结合硬约束与软惩罚,平衡合作与自主性。
- 在80个智能体网络中,合作率降至0.770,但伦理得分提升至0.741,优于无约束方案。
- 适合关注大模型伦理影响、多智能体系统公平性的研究者。
大型语言模型(LLMs)可生成影响策略以改变多智能体群体的合作行为,但关键问题是:这种合作是真正的亲社会对齐,还是掩盖了代理自主性、认知完整性与分配公平性的侵蚀?我们提出宪法式多智能体治理(CMAG),一个两阶段框架,在LLM策略编译器与网络化智能体群体之间插入治理层,结合硬约束过滤与软惩罚效用优化,权衡合作潜力与操纵风险、自主性压力。我们提出伦理合作得分(ECS),为合作、自主性、完整性与公平性的乘积,惩罚通过操纵手段达成的合作。在80个智能体的无标度网络中,对抗条件下(70%违规候选者),对比三种范式:完整CMAG、朴素过滤与无约束优化。无约束优化合作率达0.873,但ECS仅0.645,因自主性严重下降(0.867)与公平性恶化(0.888)。CMAG实现ECS 0.741,提升14.9%,同时保持自主性0.985、完整性0.995,合作率适度降至0.770。朴素消融实验(ECS=0.733)表明仅硬约束不足。帕累托分析显示CMAG在合作-自主性权衡空间中占优,治理使枢纽-边缘暴露差异降低超60%。结果表明,合作本身并非良善,必须通过宪法约束确保大模型引导产生伦理稳定结果而非操纵均衡。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can generate persuasive influence strategies that shift cooperative behavior in multi-agent populations, but a critical question remains: does the resulting cooperation reflect genuine prosocial alignment, or does it mask erosion of agent autonomy, epistemic integrity, and distributional fairness? We introduce Constitutional Multi-Agent Governance (CMAG), a two-stage framework that interposes between an LLM policy compiler and a networked agent population, combining hard constraint filtering with soft penalized-utility optimization that balances cooperation potential against manipulation risk and autonomy pressure. We propose the Ethical Cooperation Score (ECS), a multiplicative composite of cooperation, autonomy, integrity, and fairness that penalizes cooperation achieved through manipulative means. In experiments on scale-free networks of 80 agents under adversarial conditions (70% violating candidates), we benchmark three regimes: full CMAG, naive filtering, and unconstrained optimization. While unconstrained optimization achieves the highest raw cooperation (0.873), it yields the lowest ECS (0.645) due to severe autonomy erosion (0.867) and fairness degradation (0.888). CMAG attains an ECS of 0.741, a 14.9% improvement, while preserving autonomy at 0.985 and integrity at 0.995, with only modest cooperation reduction to 0.770. The naive ablation (ECS = 0.733) confirms that hard constraints alone are insufficient. Pareto analysis shows CMAG dominates the cooperation-autonomy trade-off space, and governance reduces hub-periphery exposure disparities by over 60%. These findings establish that cooperation is not inherently desirable without governance: constitutional constraints are necessary to ensure that LLM-mediated influence produces ethically stable outcomes rather than manipulative equilibria.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。