用上下文学习让智能体自发协作,无需预设规则。
Multi-agent cooperation through in-context co-player inference
- 利用序列模型的上下文学习能力,动态推断对手策略。
- 在多样化对手环境中训练,自然涌现合作行为。
- 无需预设学习规则,适合大规模多智能体系统。
多智能体强化学习中,如何实现自利智能体间的协作仍是核心挑战。已有方法依赖对对手学习规则的硬编码假设,或强制区分快速更新的‘普通学习者’与观察其更新的‘元学习者’。本文证明,序列模型的上下文学习能力可使智能体无需预设假设或显式时间尺度分离,便具备对手学习意识。在多样化的对手分布上训练序列模型智能体,能自然产生上下文最优响应策略,相当于在单个回合内快速执行学习算法。我们发现先前研究中的合作机制——因易受勒索而引发相互塑造——在此设置中自然出现:上下文适应使智能体易受勒索,由此产生的相互压力促使双方调整对方的上下文学习动态,最终演化为合作行为。结果表明,结合对手多样性与标准去中心化序列模型强化学习,为规模化学习协作行为提供可行路径。
原文摘要 · Abstract (English)
Achieving cooperation among self-interested agents remains a fundamental challenge in multi-agent reinforcement learning. Recent work showed that mutual cooperation can be induced between "learning-aware" agents that account for and shape the learning dynamics of their co-players. However, existing approaches typically rely on hardcoded, often inconsistent, assumptions about co-player learning rules or enforce a strict separation between "naive learners" updating on fast timescales and "meta-learners" observing these updates. Here, we demonstrate that the in-context learning capabilities of sequence models allow for co-player learning awareness without requiring hardcoded assumptions or explicit timescale separation. We show that training sequence model agents against a diverse distribution of co-players naturally induces in-context best-response strategies, effectively functioning as learning algorithms on the fast intra-episode timescale. We find that the cooperative mechanism identified in prior work-where vulnerability to extortion drives mutual shaping-emerges naturally in this setting: in-context adaptation renders agents vulnerable to extortion, and the resulting mutual pressure to shape the opponent's in-context learning dynamics resolves into the learning of cooperative behavior. Our results suggest that standard decentralized reinforcement learning on sequence models combined with co-player diversity provides a scalable path to learning cooperative behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。