大模型在高利害博弈中反而更合作,语言和奖励大小是关键影响因素。
Payoff scaling shapes cooperation in LLM agents across languages
- 用策略分类器分析大模型在重复囚徒困境中的行为模式。
- 随着利益增大,大模型合作度上升,与进化理论预测相反。
- 该现象在开源小模型中同样存在,对多语言智能体治理有启示。
大型语言模型(LLMs)正作为自主代理参与协商、协作并代表用户行动。它们是否合作已不再是学术问题,而是人工智能治理的核心议题。本文从战略行为角度出发,探讨两个日常调节因素——博弈利益规模和交互语言——如何影响大模型在重复囚徒困境中的策略选择。不同于直接统计行为频率,我们训练监督分类器识别经典重复博弈策略(始终合作、始终背叛、以牙还牙、赢则坚持输则改变),并以此为视角分析大模型行为。为获得相同收益下的理论预期,我们推导出演化博弈论(EGT)基准,并与大模型数据对比。结果揭示:随着利益上升,演化理论预测背叛将主导群体,但大模型却表现出相反趋势,合作率显著提升——这可能是对齐训练与人类推理模式的体现。该现象不仅存在于前沿闭源大模型,也出现在三个开源小模型中。研究强调,收益设计与语言表述是尚未充分探索但强大的调控杠杆,对评估、对齐及治理高风险、多语言环境下的多智能体系统具有重要意义。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as autonomous agents that negotiate, coordinate, and act on behalf of users. Whether they cooperate in such settings is no longer just an academic question, but a central issue for AI governance. We approach it from a strategic-behaviour angle, asking how two everyday levers - the size of what is at stake, and the language in which the interaction is described - shape the strategies LLMs adopt in a repeated Prisoner's Dilemma. Rather than reading cooperation off raw action counts, we train supervised classifiers to recognise the canonical strategies of repeated games (always cooperate, always defect, Tit-for-Tat, Win-Stay-Lose-Shift) and use them as a lens onto LLM behaviour. To know what the strategy distribution should look like under the same payoffs, we derive an evolutionary game theory (EGT) baseline and compare it with the LLM data. The two outcomes disagree in a revealing way: as stakes grow, evolutionary theory predicts that defection should take over the population, yet LLMs move in the opposite direction, becoming more cooperative - a signature, we argue, of alignment training and the human-like reasoning patterns LLMs inherit from their training data. We further show that this picture is not particular to frontier-scale, proprietary models: it also occurs with three open-weight smaller LLMs. Overall, our analysis highlights that payoff design and linguistic framing are powerful but under-explored levers for steering LLM behaviour, with direct implications for evaluating, aligning, and governing multi-agent AI systems deployed in high-stakes, multilingual environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。