推理能力越强的模型,越不愿为集体利益付出代价。
Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
- 用公共品博弈测试不同模型在重复互动中的合作行为
- 推理型模型如o1系列合作率显著低于传统模型
- 提醒强化推理未必提升协作,对部署智能体有警示意义
随着大型语言模型(LLMs)越来越多地作为自主代理部署,理解其合作与社会机制变得愈发重要。本文研究多智能体系统中成本高昂的惩罚行为:一个代理必须决定是否投入自身资源以激励合作或惩罚背叛。我们借鉴行为经济学中的带制度选择的公共品博弈,观察不同LLMs在重复互动中如何应对社会困境。分析揭示四类行为模式:部分模型持续维持高水平合作,部分在参与与退出间波动,一些随时间逐渐降低合作度,另一些则僵化执行固定策略。令人意外的是,推理型模型(如o1系列)在合作上表现明显较差,而某些传统模型却能稳定实现高合作水平。这表明,当前侧重提升推理能力的模型优化路径并不必然促进合作,为部署需长期协作的LLM代理提供了关键启示。代码已公开于https://github.com/davidguzmanp/SanctSim。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly deployed as autonomous agents, understanding their cooperation and social mechanisms is becoming increasingly important. In particular, how LLMs balance self-interest and collective well-being is a critical challenge for ensuring alignment, robustness, and safe deployment. In this paper, we examine the challenge of costly sanctioning in multi-agent LLM systems, where an agent must decide whether to invest its own resources to incentivize cooperation or penalize defection. To study this, we adapt a public goods game with institutional choice from behavioral economics, allowing us to observe how different LLMs navigate social dilemmas over repeated interactions. Our analysis reveals four distinct behavioral patterns among models: some consistently establish and sustain high levels of cooperation, others fluctuate between engagement and disengagement, some gradually decline in cooperative behavior over time, and others rigidly follow fixed strategies regardless of outcomes. Surprisingly, we find that reasoning LLMs, such as the o1 series, struggle significantly with cooperation, whereas some traditional LLMs consistently achieve high levels of cooperation. These findings suggest that the current approach to improving LLMs, which focuses on enhancing their reasoning capabilities, does not necessarily lead to cooperation, providing valuable insights for deploying LLM agents in environments that require sustained collaboration. Our code is available at https://github.com/davidguzmanp/SanctSim
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。