arXiv:2409.01646cs.RO2024-09被引 12

用时空对比学习提升机器人在鸟瞰图下的自主导航能力

BEVNav: Robot Autonomous Navigation Via Spatial-Temporal Contrastive Learning in Bird's-Eye View

  • 通过点云的时空对比学习,自动构建鸟瞰图特征表示
  • 在密集人流环境下优于当前最优方法,导航成功率显著提升
  • 适合需要无地图自主导航的机器人研发与应用

在无地图环境中,基于目标的移动机器人导航需要可靠的环境状态表征以支持稳定决策。受点云中鸟瞰图(BEV)在视觉感知中的优良特性启发,本文提出一种名为BEVNav的新导航方法。该方法采用深度强化学习学习BEV表征,提升决策可靠性。首先,提出一种自监督的时空对比学习方法:空间上,对点云的两个随机增强视图进行相互预测,强化空间特征;时间上,将当前观测与连续帧的动作结合,预测未来特征,建立观测变化与动作之间的关联,捕捉时序线索。随后,将该时空对比学习嵌入Soft Actor-Critic强化学习框架,获得更优的导航策略。大量实验表明,BEVNav在行人密集环境中表现出强鲁棒性,在多个基准测试中超越现有先进方法。

原文摘要 · Abstract (English)

Goal-driven mobile robot navigation in map-less environments requires effective state representations for reliable decision-making. Inspired by the favorable properties of Bird's-Eye View (BEV) in point clouds for visual perception, this paper introduces a novel navigation approach named BEVNav. It employs deep reinforcement learning to learn BEV representations and enhance decision-making reliability. First, we propose a self-supervised spatial-temporal contrastive learning approach to learn BEV representations. Spatially, two randomly augmented views from a point cloud predict each other, enhancing spatial features. Temporally, we combine the current observation with consecutive frames' actions to predict future features, establishing the relationship between observation transitions and actions to capture temporal cues. Then, incorporating this spatial-temporal contrastive learning in the Soft Actor-Critic reinforcement learning framework, our BEVNav offers a superior navigation policy. Extensive experiments demonstrate BEVNav's robustness in environments with dense pedestrians, outperforming state-of-the-art methods across multiple benchmarks. \rev{The code will be made publicly available at https://github.com/LanrenzzzZ/BEVNav.

自主导航鸟瞰图对比学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。