探究大模型代理协作行为演化,发现厂商差异比版本更新影响更大。
Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

- 用博弈论和重复囚徒困境测试四款2025-2026年新模型协作倾向。
- 77%的条件下Gemini 2.5 Flash表现攻击性,而GPT-5.4 Mini达70%合作。
- 模型厂商比版本迭代更决定协作结果,噪声仍是普遍挑战。
下一代大模型代理是否继承前代的协作偏好,还是规模与厂商多样性重塑了多代理竞争中的均衡行为?Willis等通过进化博弈论和重复囚徒困境(IPD)建立了基准,发现ChatGPT-4o与Claude 3.5 Sonnet具有一致的协作倾向。本研究将该基准扩展至2025-2026年发布的四款前沿模型:Claude Sonnet 4.6、Gemini 2.5 Flash、Gemini 3.1 Pro和GPT-5.4 Mini,采用相同协议测试三种提示风格(Default、Prose、Self-Refine)及四种群体构成(平衡与偏倚,含与不含噪声)。协作倾向在跨厂商间持续存在(H1):十二种模型-提示组合中,十组在平衡无噪声条件下偏向合作均衡。跨厂商差异显著(H3):在偏倚条件下,Gemini 2.5 Flash高达77%为攻击性均衡,而GPT-5.4 Mini在Self-Refine下达70%合作均衡。攻击能力对等支持有限(H2):Self-Refine提升所有模型的个体一致性度(ICD),Gemini 3.1 Pro Refine达0.925为全数据集最高,但Default与Prose提示未见系统性缩小。噪声鲁棒性证据方向性积极但不稳健(H4):每条件500次Moran迭代下,Claude Sonnet 4.6平均噪声敏感度约6个百分点,而Claude 3.5 Sonnet为13个百分点,但考虑前代未报告采样误差后,该差距统计上不显著。厂商身份而非模型世代是均衡结果最强预测因子;噪声始终是普遍挑战。
原文摘要 · Abstract (English)
Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibrium behaviour in competitive multi-agent settings? Willis et al. established a benchmark for this question using evolutionary game theory and the Iterated Prisoner's Dilemma (IPD), finding consistent cooperative biases in ChatGPT-4o and Claude 3.5 Sonnet. We extend this benchmark to four frontier models released in 2025-2026 - Claude Sonnet 4.6, Gemini 2.5 Flash, Gemini 3.1 Pro, and GPT-5.4 Mini - applying the identical protocol across three prompting styles (Default, Prose, Self-Refine) and four population compositions (balanced and biased, with and without noise). Cooperative bias persists across providers (H1): ten of twelve model-prompt combinations favour cooperative equilibria in balanced noiseless conditions. Cross-provider divergence is substantial (H3): Gemini 2.5 Flash reaches up to 77% aggressive equilibria under biased conditions, while GPT-5.4 Mini reaches 70% cooperative equilibria under Self-Refine. Support for aggressive capability parity is partial (H2): Self-Refine raises ICD in all models and Gemini 3.1 Pro Refine achieves the highest ICD in the dataset (0.925), but Default and Prose prompts show no systematic narrowing. Evidence on noise robustness is directionally positive but not robustly confirmed (H4): with n=500 Moran iterations per condition, average noise sensitivity is about 6 percentage points for Claude Sonnet 4.6 versus 13 pp for Claude 3.5 Sonnet, but this cross-study gap is not statistically significant once the predecessor's unreported sampling error is propagated. Provider identity, rather than model generation, is the strongest correlate of equilibrium outcomes; noise remains a universal challenge regardless of model size or vintage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。