arXiv:2512.00352cs.LGstat.ML2025-12NeurIPS被引 1

提出首个样本高效且鲁棒的离线双人零和博弈算法

Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement Learning

  • 基于乐观鲁棒值迭代与数据驱动惩罚项,提升离线策略稳定性
  • 在部分覆盖条件下实现近似最优样本复杂度,理论证明最优性
  • 适合研究离线强化学习鲁棒性与多智能体博弈的学者

多智能体强化学习(MARL)研究多个智能体在共享动态环境中的独立决策。由于环境不确定性,MARL策略必须具备鲁棒性以应对仿真到现实的差距。本文聚焦于离线设置下的两玩家零和马尔可夫博弈(TZMGs),特别是表格式鲁棒TZMGs(RTZMGs)。我们提出一种基于模型的算法(RTZ-VI-LCB),结合乐观鲁棒值迭代与数据驱动的伯恩斯坦风格惩罚项,用于鲁棒值估计。该算法考虑历史数据集中的分布偏移,在部分覆盖和环境不确定性下建立了近似最优的样本复杂度保证。通过信息论下界分析,验证了算法样本复杂度的紧致性,其最优性同时涵盖状态空间与动作空间。据我们所知,RTZ-VI-LCB是首个达到此最优性的方法,为离线鲁棒两玩家零和博弈设立了新基准,并通过实验验证了其有效性。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL), as a thriving field, explores how multiple agents independently make decisions in a shared dynamic environment. Due to environmental uncertainties, policies in MARL must remain robust to tackle the sim-to-real gap. We focus on robust two-player zero-sum Markov games (TZMGs) in offline settings, specifically on tabular robust TZMGs (RTZMGs). We propose a model-based algorithm (\textit{RTZ-VI-LCB}) for offline RTZMGs, which is optimistic robust value iteration combined with a data-driven Bernstein-style penalty term for robust value estimation. By accounting for distribution shifts in the historical dataset, the proposed algorithm establishes near-optimal sample complexity guarantees under partial coverage and environmental uncertainty. An information-theoretic lower bound is developed to confirm the tightness of our algorithm's sample complexity, which is optimal regarding both state and action spaces. To the best of our knowledge, RTZ-VI-LCB is the first to attain this optimality, sets a new benchmark for offline RTZMGs, and is validated experimentally.

强化学习多智能体离线学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。