arXiv:2605.03125cs.LG2026-05

突破多智能体强化学习在大规模状态空间下的数据效率瓶颈。

Taming the Curses of Multiagency in Robust Markov Games with Large State Space through Linear Function Approximation

  • 基于线性函数近似,设计新型算法应对环境不确定性。
  • 在生成模型与在线交互两种场景下均实现样本高效,打破多智能体诅咒。
  • 首次解决大规模状态空间下鲁棒马尔可夫博弈的样本复杂度问题。

多智能体强化学习虽潜力巨大,但受环境不确定性影响,鲁棒性面临挑战。分布鲁棒马尔可夫博弈(RMGs)通过优化不确定性集内最坏情况性能来提升鲁棒性。然而,数据效率同样是关键目标:随着智能体数量增加,状态与动作空间呈指数级增长,导致多智能体诅咒。现有可证明数据高效的RMG算法仅适用于有限状态与动作空间的表格型设置,难以处理大规模或无限状态空间。现有非表格型工作仅针对特定类别的RMG,依赖极小值假设且仍存在多智能体诅咒问题。本文研究一般性RMG结合线性函数近似(LFA),针对由总变差距离定义的不确定性集,提出可证明数据高效的算法,在生成模型设置与新提出的在线交互设置中均打破多智能体诅咒。据我们所知,这是首个在大规模(可能无限)状态空间下,对任意不确定性集构造都实现样本复杂度上突破多智能体诅咒的成果。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) holds great potential but faces robustness challenges due to environmental uncertainty. To address this, distributionally robust Markov games (RMGs) optimize worst-case performance when the environment deviates from the nominal model within a uncertainty set. Beyond robustness, an equally urgent goal for MARL is data efficiency -- sampling from vast state and action spaces that grow exponentially with the number of agents potentially leads to the curse of multiagency. However, current provably data-efficient algorithms for RMGs are limited to tabular settings with finite state and action spaces, which are only computationally manageable for small-scale problems, leaving RMGs with large-scale (or infinite) state spaces largely unexplored. The only existing work beyond tabular settings focuses on linear function approximation (LFA) for a restrictive class of RMGs using vanish minimal value assumption and still suffers from sample complexity with the curse of multiagency. In this work, we focuses on general RMGs with LFA. For uncertainty sets defined by total variation distance, we develop provably data-efficient algorithms that break the curse of multiagency in both the generative model setting and a newly proposed online interactive setting. To our knowledge, our results are the first to break the curse of multiagency of sample complexity for RMGs with large (possibly infinite) state spaces, regardless of the uncertainty set construction.

多智能体强化学习鲁棒性函数逼近

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。