arXiv:2510.15679cs.RO2025-10被引 1

用注意力机制实现大场景高效探索,自动构建可扩展地图。

HEADER: Hierarchical Robot Exploration via Attention-Based Deep Reinforcement Learning with Expert-Guided Reward

  • 基于注意力网络与分层图结构,动态构建全局地图。
  • 在仿真和真实场景中探索效率提升最高达20%。
  • 无需调参的奖励机制,适合复杂大环境机器人探索。

本工作在环境规模与探索效率上推进了基于学习的自主机器人探索方法。我们提出HEADER,一种基于注意力机制的强化学习方法,结合分层图结构以实现大规模环境中的高效探索。HEADER沿用传统方法构建机器人信念/地图的分层表示,进一步设计了一种新型基于社区的算法来构建和更新全局图,该图保持完全增量式、形状自适应,并具有线性复杂度。基于注意力网络的规划器在局部范围内精细推理邻近信念,同时粗粒度利用远距离信息,实现考虑多尺度空间依赖性的最优视角决策。除新颖的地图表示外,我们引入一种无参数的特权奖励,显著提升模型性能并生成接近最优的探索行为,避免了手工奖励设计带来的训练目标偏差。在模拟的复杂大规模探索场景中,HEADER表现出优于大多数现有学习与非学习方法的可扩展性,探索效率较最先进基线提升高达20%。我们还在硬件上部署HEADER,验证其在复杂真实场景中的有效性,包括一个300m×230m的校园环境。

原文摘要 · Abstract (English)

This work pushes the boundaries of learning-based methods in autonomous robot exploration in terms of environmental scale and exploration efficiency. We present HEADER, an attention-based reinforcement learning approach with hierarchical graphs for efficient exploration in large-scale environments. HEADER follows existing conventional methods to construct hierarchical representations for the robot belief/map, but further designs a novel community-based algorithm to construct and update a global graph, which remains fully incremental, shape-adaptive, and operates with linear complexity. Building upon attention-based networks, our planner finely reasons about the nearby belief within the local range while coarsely leveraging distant information at the global scale, enabling next-best-viewpoint decisions that consider multi-scale spatial dependencies. Beyond novel map representation, we introduce a parameter-free privileged reward that significantly improves model performance and produces near-optimal exploration behaviors, by avoiding training objective bias caused by handcrafted reward shaping. In simulated challenging, large-scale exploration scenarios, HEADER demonstrates better scalability than most existing learning and non-learning methods, while achieving a significant improvement in exploration efficiency (up to 20%) over state-of-the-art baselines. We also deploy HEADER on hardware and validate it in complex, large-scale real-life scenarios, including a 300m*230m campus environment.

机器人探索强化学习注意力机制大场景导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。