arXiv:2508.02948cs.LGcs.MA2025-08中稿 · ICLR被引 5

在线学习让多智能体系统在不确定环境中更鲁棒,无需预训练数据。

Sample-Efficient Distributionally Robust Multi-Agent Reinforcement Learning via Online Interaction

  • 基于在线交互设计新算法MORNAVI,直接从环境反馈中学习。
  • 理论证明算法在总变差与KL散度下均实现低遗憾并找到最优鲁棒策略。
  • 适合追求真实场景鲁棒性的多智能体系统研究者使用。

训练良好的多智能体系统在部署时可能因训练与实际环境间的模型不匹配而失效,此类不匹配源于环境噪声或对抗攻击等不确定性。分布鲁棒马尔可夫博弈(DRMG)通过优化最坏情况下的性能来增强系统韧性。然而,现有方法依赖模拟器或大规模离线数据集,这些资源往往不可得。本文首次研究了DRMG中的在线学习,即智能体在无先验数据的情况下直接从环境交互中学习。我们提出多玩家乐观鲁棒纳什值迭代(MORNAVI)算法,并首次为该设置提供了可证明的保证。理论分析表明,该算法在总变差和KL散度度量的不确定性集合下均实现低遗憾,并高效找到最优鲁棒策略。这些结果为构建真正鲁棒的多智能体系统开辟了一条新路径。

原文摘要 · Abstract (English)

Well-trained multi-agent systems can fail when deployed in real-world environments due to model mismatches between the training and deployment environments, caused by environment uncertainties including noise or adversarial attacks. Distributionally Robust Markov Games (DRMGs) enhance system resilience by optimizing for worst-case performance over a defined set of environmental uncertainties. However, current methods are limited by their dependence on simulators or large offline datasets, which are often unavailable. This paper pioneers the study of online learning in DRMGs, where agents learn directly from environmental interactions without prior data. We introduce the Multiplayer Optimistic Robust Nash Value Iteration (MORNAVI) algorithm and provide the first provable guarantees for this setting. Our theoretical analysis demonstrates that the algorithm achieves low regret and efficiently finds the optimal robust policy for uncertainty sets measured by Total Variation divergence and Kullback-Leibler divergence. These results establish a new, practical path toward developing truly robust multi-agent systems.

多智能体鲁棒学习在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。