arXiv:2602.22810cs.LG2026-02被引 1

提出线性马尔可夫博弈下高效多智能体模仿学习方法,突破传统样本瓶颈。

Multi-agent imitation learning with function approximation: Linear Markov games and beyond

  • 利用特征层面的集中性系数替代状态动作级,降低复杂度。
  • 设计首个计算高效的交互式算法,样本复杂度仅依赖特征维数d。
  • 基于理论成果的深度算法在井字棋和连珠棋中显著优于行为克隆。

本文首次对线性马尔可夫博弈中的多智能体模仿学习(MAIL)进行理论分析,其中转移动态与各智能体奖励函数均关于给定特征呈线性。通过利用该结构,可将原有的状态-动作级“所有策略偏离集中性系数”替换为特征级集中性系数,当特征能有效反映状态相似性时,该系数可远小于原版本。为进一步避免集中性系数依赖,转向交互设置,提出首个计算高效的交互式邮件算法,其样本复杂度仅与特征映射维度d相关。基于上述理论成果,设计了一种深度交互式邮件算法,在井字棋和连珠棋等游戏中明显优于行为克隆(BC)。

原文摘要 · Abstract (English)

In this work, we present the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games where both the transition dynamics and each agent's reward function are linear in some given features. We demonstrate that by leveraging this structure, it is possible to replace the state-action level "all policy deviation concentrability coefficient" (Freihaut et al., arXiv:2510.09325) with a concentrability coefficient defined at the feature level which can be much smaller than the state-action analog when the features are informative about states' similarity. Furthermore, to circumvent the need for any concentrability coefficient, we turn to the interactive setting. We provide the first, computationally efficient, interactive MAIL algorithm for linear Markov games and show that its sample complexity depends only on the dimension of the feature map $d$. Building on these theoretical findings, we propose a deep MAIL interactive algorithm which clearly outperforms BC on games such as Tic-Tac-Toe and Connect4.

多智能体模仿学习线性马尔可夫博弈样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。