arXiv:2510.09325cs.LGcs.AI2025-10被引 2

提出首个近似最优的交互式多智能体模仿学习算法,突破样本效率瓶颈。

Rate optimal learning of equilibria from data

  • 结合无奖赏强化学习与交互式模仿学习,设计新框架MAIL-WARM
  • 样本复杂度从O(ε⁻⁸)降至O(ε⁻²),逼近理论下限
  • 揭示行为克隆在复杂环境中的局限性,适合多智能体系统研究者

我们通过刻画非交互式多智能体模仿学习(MAIL)的理论极限,填补了该领域的开放问题,并提出了首个具有近似最优样本复杂度的交互式算法。在非交互设置下,我们证明了统计下界,指出所有策略偏差集中系数是根本复杂度度量,且行为克隆(BC)达到率最优。在交互设置中,我们引入一个结合无奖赏强化学习与交互式MAIL的框架,并以MAIL-WARM算法为例实现。其样本复杂度由先前最优的O(ε⁻⁸)提升至O(ε⁻²),与我们的下界对ε的依赖一致。最后,我们在网格世界等环境中提供数值结果,验证理论并展示行为克隆无法学习的情形。

原文摘要 · Abstract (English)

We close open theoretical gaps in Multi-Agent Imitation Learning (MAIL) by characterizing the limits of non-interactive MAIL and presenting the first interactive algorithm with near-optimal sample complexity. In the non-interactive setting, we prove a statistical lower bound that identifies the all-policy deviation concentrability coefficient as the fundamental complexity measure, and we show that Behavior Cloning (BC) is rate-optimal. For the interactive setting, we introduce a framework that combines reward-free reinforcement learning with interactive MAIL and instantiate it with an algorithm, MAIL-WARM. It improves the best previously known sample complexity from $\mathcal{O}(\varepsilon^{-8})$ to $\mathcal{O}(\varepsilon^{-2}),$ matching the dependence on $\varepsilon$ implied by our lower bound. Finally, we provide numerical results that support our theory and illustrate, in environments such as grid worlds, where Behavior Cloning fails to learn.

多智能体模仿学习样本效率强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。