arXiv:2606.10832cs.RO2026-06

让四足机器人仅凭初始目标实现全程自主导航

GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation

论文配图:GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation
图 1 · 摘自论文原文
  • 用本体感知历史构建空间锚点,维持长期方向记忆
  • 无需外部目标更新,在复杂环境成功避障通过
  • 适合追求完全端到端自主的机器人研究者

基于学习的腿部机器人视觉导航通常依赖分层状态估计持续提供目标更新,带来额外传感与计算开销,并违背全端到端自主。在部分可观测环境下,策略易陷入短视行为,困于死胡同或复杂结构。为此,本文研究仅在任务开始时提供一次目标的导航设定,要求机器人依靠内在空间记忆运行,无需后续目标更新。提出全新端到端强化学习框架GUIDE,通过多频次本体感知历史提取自运动表征,构建持久长时程空间上下文;同时利用原始深度流感知局部环境几何。在仿真与真实四足机器人场景中验证,结果表明GUIDE能学习可靠的自运动与方向意识,使全端到端策略在无后续目标引导或先验地图条件下,安全穿越密集障碍物与结构化迷宫。

原文摘要 · Abstract (English)

Learning-based visual navigation for legged robots typically relies on continuous goal updates from hierarchical state estimation to provide a persistent directional reference. This reliance incurs additional sensory and computational overhead and deviates from fully end-to-end mobile autonomy. Furthermore, under partial observability, policies are prone to learn myopic behaviors, easily becoming trapped in dead ends and complex structural layouts. To address these limitations, we investigate a goal-initialized navigation setting, where the target is provided only once at the beginning of an episode, requiring the robot to operate based on intrinsic spatial memory without subsequent goal updates from external modules. In this work, we propose GUIDE, a fully end-to-end reinforcement learning framework designed to cultivate internal directional awareness. Specifically, GUIDE incorporates a spatial anchor predictor that leverages multi-frequency proprioceptive history to extract egomotion representations, thereby maintaining a persistent long-horizon spatial context for navigation. Concurrently, it utilizes raw depth streams to perceive local environmental geometry. We evaluate the proposed framework across both simulation and real-world scenarios on a quadruped robot. Experiments show that GUIDE learns reliable egomotion and directional awareness, enabling a fully end-to-end deployed policy to safely navigate through dense clutter and structured mazes without subsequent goal guidance or prior maps.

视觉导航四足机器人端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。