比较了多智能体系统中规则自生与外优化的优劣,发现进化更优但依赖激励机制。
Internal vs. External: Comparing Deliberation and Evolution for Multi-Agent Constitutional Design
- 用演化算法和内部协商两种方式设计规则,对比其在三种社会环境中的表现。
- 演化在集体行动场景中显著优于协商(p<0.01),但在双边交易中无优势。
- 当激励系数降低至0.75时,演化反而导致合作失败,凸显机制敏感性。
多智能体人工智能系统需要行为规范,但尚未明确这些规则应通过智能体自我治理内生生成,还是通过外部优化发现。我们首次在三个社会环境中对内部协商与外部演化进行了受控比较:一个协调网格世界、一个重复公共品博弈和一个双边交易市场。在180次模拟运行中,演化在集体行动场景中显著优于协商(p < 0.01),而在双边交易中两者均未提升结果。乘数消融实验显示,当池乘数(m = 0.75)降低时,演化的规则导致破坏价值的合作,成为表现最差的方法。值得注意的是,在三十次试验中,所有协商过程均未提出惩罚机制——而该机制是演化可靠发现的维持合作的关键策略,表明外部优化在关键峰点上占优,而内部自洽则以结构响应性为代价换取稳健性。
原文摘要 · Abstract (English)
Multi-agent AI systems need behavioral constitutions, but it is unresolved whether such rules should emerge internally through agent self-governance or be discovered externally through optimization. We present the first controlled comparison of internal deliberation and external evolution across three social environments: a coordination grid-world, an iterated public goods game, and a bilateral trading market. Across 180 simulation runs, evolution significantly outperforms deliberation in collective-action settings (p < 0.01), while neither method improves outcomes in bilateral trading. A multiplier ablation reveals that evolution's advantage inverts when incentives shift: at pool multiplier (m = 0.75) the evolved constitution forces value-destroying cooperation and becomes the worst-performing method. Notably, no deliberation run across thirty trials ever proposed punishment -- the canonical cooperation-sustaining mechanism evolution reliably discovers -- suggesting external optimization wins on peaks while internal self-governance trades peaks for structural responsiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。