arXiv:2505.12811cs.MAcs.AI2025-05中稿 · AAMAS 2025

动态调整智能体视野范围,提升多智能体强化学习性能。

Dynamic Sight Range Selection in Multi-Agent Reinforcement Learning

  • 用置信度上界算法动态调节视野范围。
  • 在三种环境中提升性能,加速训练过程。
  • 无需全局信息,适合实际复杂场景。

多智能体强化学习常面临视野范围困境:信息过少或过多。本文提出动态视野选择(DSR)方法,利用置信度上界(UCB)算法在训练中动态调整视野范围。实验表明,DSR在三种典型环境(层级觅食LBF、多机器人仓库RWARE、星际争霸多智能体挑战SMAC)中均表现更优;且对QMIX、MAPPO等多种MARL算法均有稳定提升;能为不同训练阶段提供合适视野,加快收敛;同时通过记录最优视野范围增强可解释性。该方法仅依赖个体视野,不需全局信息或通信机制,具有实用性与普适性。

原文摘要 · Abstract (English)

Multi-agent reinforcement Learning (MARL) is often challenged by the sight range dilemma, where agents either receive insufficient or excessive information from their environment. In this paper, we propose a novel method, called Dynamic Sight Range Selection (DSR), to address this issue. DSR utilizes an Upper Confidence Bound (UCB) algorithm and dynamically adjusts the sight range during training. Experiment results show several advantages of using DSR. First, we demonstrate using DSR achieves better performance in three common MARL environments, including Level-Based Foraging (LBF), Multi-Robot Warehouse (RWARE), and StarCraft Multi-Agent Challenge (SMAC). Second, our results show that DSR consistently improves performance across multiple MARL algorithms, including QMIX and MAPPO. Third, DSR offers suitable sight ranges for different training steps, thereby accelerating the training process. Finally, DSR provides additional interpretability by indicating the optimal sight range used during training. Unlike existing methods that rely on global information or communication mechanisms, our approach operates solely based on the individual sight ranges of agents. This approach offers a practical and efficient solution to the sight range dilemma, making it broadly applicable to real-world complex environments.

多智能体强化学习动态调整视野范围

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。