LLM在博弈中易背叛,新基准测试发现签约和中介最促合作。
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

- 设计四类机制:重复游戏、声誉系统、第三方调解、合同约束
- 签约与调解使强推理LLM在博弈中合作率显著提升,重复游戏效果差
- 机制在追求自身利益的演化压力下更有效,适合安全协作研究者
随着大语言模型(LLM)在多智能体系统中的应用日益广泛,其与目标导向智能体的有效且安全互动变得愈发重要。然而,近期研究表明,具备更强推理能力的LLM在囚徒困境、公共品等混合动机博弈中反而表现出更低的合作意愿。我们实验发现,即使启用推理功能,当前主流模型在单轮社会困境中仍持续选择背叛。为应对这一安全挑战,本文首次对旨在实现理性智能体间均衡合作的博弈论机制进行对比研究。在四种测试不同合作韧性维度的社会困境中,评估了四类机制:(1)多轮重复博弈,(2)声誉系统,(3)委托决策给第三方调解者,(4)基于结果条件支付的合同协议。结果显示,合同机制与调解机制在促成强能力LLM之间的合作方面最为有效;而重复博弈带来的合作在同伴不一致时急剧下降。此外,我们发现这些机制在追求个体收益最大化的演化压力下表现更佳。
原文摘要 · Abstract (English)
It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite trend: LLMs with stronger reasoning capabilities behave _less_ cooperatively in mixed-motive games such as the prisoner's dilemma and public goods settings. Indeed, our experiments show that recent models -- with or without reasoning enabled -- consistently defect in single-shot social dilemmas. To tackle this safety concern, we present the first comparative study of game-theoretic mechanisms designed to enable cooperative outcomes between rational agents _in equilibrium_. Across four social dilemmas testing distinct components of robust cooperation, we evaluate four families of mechanisms: (1) repeating the game for many rounds, (2) reputation systems, (3) third-party mediators to delegate decision making to, and (4) contract agreements for outcome-conditional payments between players. Among our findings, we establish that contracting and mediation are most effective in achieving cooperative outcomes between capable LLM models, and that repetition-induced cooperation deteriorates drastically when co-players vary. Moreover, we demonstrate that the mechanisms become _more effective_ under evolutionary pressures to maximize individual payoffs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。