arXiv:2409.19363cs.MAcs.AI2024-09AAAI被引 2

为多智能体博弈设计策略表示方法,提升模仿学习效果

Learning Strategy Representation for Imitation Learning in Multi-Agent Games

  • 通过无标识轨迹学习策略表示,无需事先知道玩家身份
  • 可识别主导策略轨迹,过滤低效数据,显著提升性能
  • 适合作为插件集成到现有模仿学习框架中

多智能体博弈的离线模仿学习数据集通常包含多种策略的玩家轨迹,需防止学习算法吸收不良行为。学习轨迹的策略表示是刻画示范者策略的有效方法。然而,现有方法常依赖玩家身份标识或强假设,不适用于多智能体场景。为此,本文提出策略表示模仿学习(STRIL)框架,能有效学习多智能体博弈中的策略表示,基于表示估计指标,并利用指标过滤次优数据。STRIL为可插拔方法,可集成至现有模仿学习算法。我们在双人乒乓球、限注德州扑克和四子棋等竞争性多智能体场景中验证其有效性。该方法成功获得策略表示与指标,准确识别主导轨迹,在各环境中显著提升现有模仿学习性能。

原文摘要 · Abstract (English)

The offline datasets for imitation learning (IL) in multi-agent games typically contain player trajectories exhibiting diverse strategies, which necessitate measures to prevent learning algorithms from acquiring undesirable behaviors. Learning representations for these trajectories is an effective approach to depicting the strategies employed by each demonstrator. However, existing learning strategies often require player identification or rely on strong assumptions, which are not appropriate for multi-agent games. Therefore, in this paper, we introduce the Strategy Representation for Imitation Learning (STRIL) framework, which (1) effectively learns strategy representations in multi-agent games, (2) estimates proposed indicators based on these representations, and (3) filters out sub-optimal data using the indicators. STRIL is a plug-in method that can be integrated into existing IL algorithms. We demonstrate the effectiveness of STRIL across competitive multi-agent scenarios, including Two-player Pong, Limit Texas Hold'em, and Connect Four. Our approach successfully acquires strategy representations and indicators, thereby identifying dominant trajectories and significantly enhancing existing IL performance across these environments.

模仿学习多智能体策略表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。