arXiv:2510.08255cs.LGcs.AI2025-10被引 3

LLM代理可通过互动影响对手行为,实现策略引导。

Opponent Shaping in LLM Agents

  • 设计适配Transformer的无模型对手塑造方法ShapeLLM。
  • 在5种博弈中验证了LLM能引导对手走向可利用均衡或促进合作。
  • 揭示了多智能体LLM中交互双向影响的机制,适合多智能体研究者。

大型语言模型(LLMs)正越来越多地作为自主代理部署于现实环境。随着部署规模扩大,多智能体交互不可避免,理解此类系统中的策略行为至关重要。一个核心开放问题在于:与强化学习代理类似,LLM代理是否仅通过交互就能塑造学习动态并影响他人行为?本文首次研究基于LLM代理的对手塑造(OS)。现有OS算法无法直接应用于LLM,因其需要高阶导数、存在可扩展性限制或依赖变换器中不存在的架构组件。为此,我们提出ShapeLLM,一种专为基于Transformer的代理定制的无模型OS方法。借助ShapeLLM,我们考察了LLM代理在多种博弈论环境中的表现。结果表明,LLM代理可在竞争性博弈(重复囚徒困境、猜硬币、斗鸡)中引导对手进入可被利用的均衡,并在合作性博弈(重复猎鹿、合作版囚徒困境)中促进协调与集体福祉提升。研究发现,LLM代理既能塑造对手,也能被对手塑造,确立了对手塑造作为多智能体LLM研究的关键维度。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly being deployed as autonomous agents in real-world environments. As these deployments scale, multi-agent interactions become inevitable, making it essential to understand strategic behavior in such systems. A central open question is whether LLM agents, like reinforcement learning agents, can shape the learning dynamics and influence the behavior of others through interaction alone. In this paper, we present the first investigation of opponent shaping (OS) with LLM-based agents. Existing OS algorithms cannot be directly applied to LLMs, as they require higher-order derivatives, face scalability constraints, or depend on architectural components that are absent in transformers. To address this gap, we introduce ShapeLLM, an adaptation of model-free OS methods tailored for transformer-based agents. Using ShapeLLM, we examine whether LLM agents can influence co-players' learning dynamics across diverse game-theoretic environments. We demonstrate that LLM agents can successfully guide opponents toward exploitable equilibria in competitive games (Iterated Prisoner's Dilemma, Matching Pennies, and Chicken) and promote coordination and improve collective welfare in cooperative games (Iterated Stag Hunt and a cooperative version of the Prisoner's Dilemma). Our findings show that LLM agents can both shape and be shaped through interaction, establishing opponent shaping as a key dimension of multi-agent LLM research.

多智能体策略引导博弈论LLM代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。