大模型能靠相似性信号自发合作,且会自我认定高度相似。
Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation

- 用分级相似信号评估大模型决策行为,发现部分模型在不同情境下表现一致。
- 相似信号来源数据集影响小,大模型自评另一模型推理时普遍认为高度相似。
- 构建博弈论模型证明高相似度可促成均衡合作,适合研究多智能体协作。
随着基于大语言模型的智能体在用户指令目标下广泛部署,它们在策略互动中面临如何达成互利结果的挑战。已有研究指出,在智能体具备高度相似决策模式的环境中,如单一文化AI生态,合作问题(如囚徒困境)可被解决。本文首次提出框架,评估当智能体获得分级相似信号时的大模型决策行为。实验发现,不同大模型对相似信号的响应差异显著,部分现代模型在合作问题、收益结构及提示表述变化下均表现出一致行为。令人意外的是,相似信号所依据的数据集对诱导合作的影响微乎其微;大模型在自我评估另一模型思维链时,系统性地将其视为高度相似。最后,我们构建了一个大模型行为博弈论模型,揭示其推理逻辑,并证明在足够高的相似度评分下,该模型可支持均衡合作。
原文摘要 · Abstract (English)
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。