多智能体强化学习让小型无人机在大场景中高效探索,兼顾视野限制与协同决策。
MARVEL: Multi-Agent Reinforcement Learning for constrained field-of-View multi-robot Exploration in Large-scale environments
- 用图注意力网络和视角融合技术,实现分布式协同探索策略。
- 在90m×90m大场景中表现优于现有方法,支持不同无人机数量和传感器配置。
- 无需重新训练即可适应新环境,实测验证了真实无人机部署可行性。
在多机器人探索中,一组移动机器人需高效测绘未知环境。尽管多数规划器假设全向传感器(如激光雷达),但小型机器人(如无人机)受限于载重,往往只能使用轻量级定向传感器(如摄像头),导致视野受限(FoV)。这使探索问题更复杂,不仅需优化机器人位置,还需控制传感器朝向。本文提出MARVEL,一种基于神经网络的框架,结合图注意力网络与新颖的前缘点与朝向特征融合技术,利用多智能体强化学习(MARL)为视野受限的机器人设计分布式协作策略。为应对视角规划带来的巨大动作空间,引入信息驱动的动作剪枝策略。MARVEL在大规模复杂室内环境中提升多机器人协同效率,且无需额外训练即可适配不同团队规模与传感器配置(如视场角和探测范围)。大量实验表明,其学习到的策略展现出有效协同行为,在多项指标上超越当前最优探索算法。我们通过实测验证了其在最大90m×90m场景下的泛化能力,并成功部署于真实无人机硬件团队。
原文摘要 · Abstract (English)
In multi-robot exploration, a team of mobile robot is tasked with efficiently mapping an unknown environments. While most exploration planners assume omnidirectional sensors like LiDAR, this is impractical for small robots such as drones, where lightweight, directional sensors like cameras may be the only option due to payload constraints. These sensors have a constrained field-of-view (FoV), which adds complexity to the exploration problem, requiring not only optimal robot positioning but also sensor orientation during movement. In this work, we propose MARVEL, a neural framework that leverages graph attention networks, together with novel frontiers and orientation features fusion technique, to develop a collaborative, decentralized policy using multi-agent reinforcement learning (MARL) for robots with constrained FoV. To handle the large action space of viewpoints planning, we further introduce a novel information-driven action pruning strategy. MARVEL improves multi-robot coordination and decision-making in challenging large-scale indoor environments, while adapting to various team sizes and sensor configurations (i.e., FoV and sensor range) without additional training. Our extensive evaluation shows that MARVEL's learned policies exhibit effective coordinated behaviors, outperforming state-of-the-art exploration planners across multiple metrics. We experimentally demonstrate MARVEL's generalizability in large-scale environments, of up to 90m by 90m, and validate its practical applicability through successful deployment on a team of real drone hardware.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。