通过建模合作偏好,让多智能体在博弈中自适应地优化策略。
Preference-based opponent shaping in differentiable games
- 引入偏好参数融入损失函数,使智能体学习时考虑对手的收益。
- 在多种可微博弈中,实现更优的奖励分配与策略收敛。
- 适合需要自适应合作或竞争的动态博弈场景,泛化性强。
多智能体博弈中的策略学习极具挑战性,因每个智能体的奖励依赖于联合策略,单纯追求自身最大收益易陷入局部最优。现有对手建模与塑造方法虽提升学习效率,但通常仅基于简单策略预测,缺乏对合作、竞争等行为偏好的建模,适用范围受限且泛化能力弱。本文提出一种基于偏好的对手塑造(PBOS)方法,通过在智能体损失函数中引入偏好参数,使其在策略更新时直接考量对手的损失函数。偏好参数与策略同步更新,使智能体能适应任意合作或竞争环境。实验验证了PBOS在多种可微博弈中的有效性,结果表明该方法能引导智能体学习合适的偏好参数,从而在多个环境中实现更优的奖励分布。
原文摘要 · Abstract (English)
Strategy learning in game environments with multi-agent is a challenging problem. Since each agent's reward is determined by the joint strategy, a greedy learning strategy that aims to maximize its own reward may fall into a local optimum. Recent studies have proposed the opponent modeling and shaping methods for game environments. These methods enhance the efficiency of strategy learning by modeling the strategies and updating processes of other agents. However, these methods often rely on simple predictions of opponent strategy changes. Due to the lack of modeling behavioral preferences such as cooperation and competition, they are usually applicable only to predefined scenarios and lack generalization capabilities. In this paper, we propose a novel Preference-based Opponent Shaping (PBOS) method to enhance the strategy learning process by shaping agents' preferences towards cooperation. We introduce the preference parameter, which is incorporated into the agent's loss function, thus allowing the agent to directly consider the opponent's loss function when updating the strategy. We update the preference parameters concurrently with strategy learning to ensure that agents can adapt to any cooperative or competitive game environment. Through a series of experiments, we verify the performance of PBOS algorithm in a variety of differentiable games. The experimental results show that the PBOS algorithm can guide the agent to learn the appropriate preference parameters, so as to achieve better reward distribution in multiple game environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。