arXiv:2602.21020cs.LGcs.GT2026-02

揭示多智能体模仿学习中策略可被利用的理论极限,提出在特定条件下可实现近似纳什均衡。

Matching Multiple Experts: On the Exploitability of Multi-Agent Imitation Learning

  • 通过反例证明完全匹配专家行为仍可能不可靠
  • 在主导策略假设下,纳什差距为 O(nε_BC/(1-γ)²)
  • 适用于有稳定策略结构的多智能体系统研究

多智能体模仿学习(MA-IL)旨在从多智能体交互的专家示范中学习最优策略。尽管已有学习策略性能的保证,但对所学策略与纳什均衡之间的距离尚无刻画。本文展示了在一般n人马尔可夫博弈中,学习低可利用性策略的不可能性和计算困难性。通过构造反例表明,即使精确匹配度量也无效,并揭示了在给定度量匹配误差下刻画纳什差距的新困难。随后,我们证明在专家均衡具有战略优势的前提下,这些挑战可被克服。具体而言,在主导策略专家均衡下,若行为克隆误差为ε_BC,可得纳什模仿差距为O(nε_BC/(1−γ)²),其中γ为折扣因子。我们进一步引入最佳响应连续性的新概念,并论证标准正则化技术隐含鼓励该性质。

原文摘要 · Abstract (English)

Multi-agent imitation learning (MA-IL) aims to learn optimal policies from expert demonstrations of interactions in multi-agent interactive domains. Despite existing guarantees on the performance of the resulting learned policies, characterizations of how far the learned polices are from a Nash equilibrium are missing for offline MA-IL. In this paper, we demonstrate impossibility and hardness results of learning low-exploitable policies in general $n$-player Markov Games. We do so by providing examples where even exact measure matching fails, and demonstrating a new hardness result on characterizing the Nash gap given a fixed measure matching error. We then show how these challenges can be overcome using strategic dominance assumptions on the expert equilibrium. Specifically, for the case of dominant strategy expert equilibria, assuming Behavioral Cloning error $ε_{\text{BC}}$, this provides a Nash imitation gap of $\mathcal{O}\left(nε_{\text{BC}}/(1-γ)^2\right)$ for a discount factor $γ$. We generalize this result with a new notion of best-response continuity, and argue that this is implicitly encouraged by standard regularization techniques.

多智能体模仿学习纳什均衡策略分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。