用视觉示范训练通用机器人导航,少数据也能实时避障。
Embodiment-Agnostic Navigation Policy Trained with Visual Demonstrations
- 用深度图和目标相对位置降低输入复杂度
- 小样本训练下实现室内外多场景实时避障导航
- 适合无特定机械结构的通用导航任务
在非结构化环境中导航对机器人而言是一项挑战。虽然强化学习有效,但常需大量数据采集且存在风险。相比之下,从专家示范中学习更高效。然而,许多现有方法依赖特定机器人形态、预设目标图像并需要大规模数据集。我们提出基于视觉示范的通用形态导航(ViDEN)框架,利用深度图降低输入维度,并依赖目标相对位置,提升对多样化环境的适应性。通过在以任务为中心且与形态无关的示范上训练扩散策略,ViDEN 能实时生成无碰撞、自适应的轨迹。在人类抓取与追踪任务上的实验表明,ViDEN 仅需少量数据即可在多种室内和室外导航场景中超越现有方法,性能优异。项目主页:https://nimicurtis.github.io/ViDEN/
原文摘要 · Abstract (English)
Learning to navigate in unstructured environments is a challenging task for robots. While reinforcement learning can be effective, it often requires extensive data collection and can pose risk. Learning from expert demonstrations, on the other hand, offers a more efficient approach. However, many existing methods rely on specific robot embodiments, pre-specified target images and require large datasets. We propose the Visual Demonstration-based Embodiment-agnostic Navigation (ViDEN) framework, a novel framework that leverages visual demonstrations to train embodiment-agnostic navigation policies. ViDEN utilizes depth images to reduce input dimensionality and relies on relative target positions, making it more adaptable to diverse environments. By training a diffusion-based policy on task-centric and embodiment-agnostic demonstrations, ViDEN can generate collision-free and adaptive trajectories in real-time. Our experiments on human reaching and tracking demonstrate that ViDEN outperforms existing methods, requiring a small amount of data and achieving superior performance in various indoor and outdoor navigation scenarios. Project website: https://nimicurtis.github.io/ViDEN/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。