arXiv:2509.04847cs.AI2025-09

语言模型在长期博弈中表现如人类般合作且灵活应变。

Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory

  • 用迭代囚徒困境测试模型合作行为,对比240种经典策略。
  • 模型表现媲美甚至超越顶尖传统策略,具备友善、可激怒、慷慨等特质。
  • 能快速识别对手策略变化,适应性接近人类,适合研究人机协作。

语言模型正越来越多地部署于交互式在线环境,从个人聊天助手到领域专用智能体,引发对其在多主体场景中合作与竞争行为的关注。现有研究多集中于孤立或短期博弈情境下的模型决策,却忽视了长期互动、人机协作及行为模式随时间演变的问题。本文通过迭代囚徒困境(IPD)这一经典框架,考察语言模型行为动态。我们让基于模型的代理与240种成熟的经典策略在类似Axelrod锦标赛的设置中对战,发现语言模型的表现达到甚至超过最知名的传统策略水平。行为分析显示,语言模型展现出强合作策略的关键特征:友善性、可激怒性和慷慨性,同时具备对对手策略变化的快速适应能力。在受控的“策略切换”实验中,模型仅需数回合即可察觉并响应对手策略转变,其适应速度可与人类相媲美甚至更优。该结果首次系统刻画了语言模型代理在长期合作中的行为特性,为未来探索其在更复杂混合人机社会环境中的角色奠定了基础。

原文摘要 · Abstract (English)

Language models are increasingly deployed in interactive online environments, from personal chat assistants to domain-specific agents, raising questions about their cooperative and competitive behavior in multi-party settings. While prior work has examined language model decision-making in isolated or short-term game-theoretic contexts, these studies often neglect long-horizon interactions, human-model collaboration, and the evolution of behavioral patterns over time. In this paper, we investigate the dynamics of language model behavior in the iterated prisoner's dilemma (IPD), a classical framework for studying cooperation and conflict. We pit model-based agents against a suite of 240 well-established classical strategies in an Axelrod-style tournament and find that language models achieve performance on par with, and in some cases exceeding, the best-known classical strategies. Behavioral analysis reveals that language models exhibit key properties associated with strong cooperative strategies - niceness, provocability, and generosity while also demonstrating rapid adaptability to changes in opponent strategy mid-game. In controlled "strategy switch" experiments, language models detect and respond to shifts within only a few rounds, rivaling or surpassing human adaptability. These results provide the first systematic characterization of long-term cooperative behaviors in language model agents, offering a foundation for future research into their role in more complex, mixed human-AI social environments.

博弈论人机协作语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。