融合语义与几何规划,实现安全可靠的视觉导航
SemGeoNav:A Safety-Guided Visual Navigation Approach with Semantic Reasoning and Geometric Planning

- 分层架构结合端到端语义推理与几何规划
- 真实场景下成功率更高,导航时间更短
- 适合需要高安全性的机器人导航任务
基于学习的视觉导航提升了语义目标到达能力,但其黑箱特性导致缺乏显式几何约束,在开放环境中常出现不可预测的避障行为。传统几何规划虽能保证安全,却难以处理高维视觉目标。为此,我们提出SemGeoNav,一种新型分层视觉导航框架,将端到端模型的高层语义推理与基于几何的方法的可靠局部规划紧密结合,实现鲁棒的图像驱动导航并显著提升避障性能。此外,引入时序轨迹平滑机制,确保机器人运动连续稳定。我们在实际环境中的Unitree Go2四足机器人上评估了SemGeoNav,结果表明其优于现有代表性方法(如ViNT和NoMaD),在成功率和导航时间上均有明显优势。
原文摘要 · Abstract (English)
Learning-based visual navigation has enhanced semantic goal-reaching capabilities. However, due to their black-box nature, purely end-to-end models often lack explicit geometric constraints, leading to unpredictable and unreliable obstacle avoidance in open environments. Conversely, traditional geometric planners ensure safety but struggle with high-dimensional visual targets. To address these limitations, we propose SemGeoNav, a novel hierarchical visual navigation framework.It tightly integrates the high-level semantic reasoning of end-to-end models with the reliable local planning ability of geometry-based methods, achieving robust image-based navigation while significantly improving obstacle avoidance. Furthermore, we introduce a temporal trajectory smoothing mechanism to ensure continuous and stable robot motion. We evaluated SemGeoNav on a Unitree Go2 quadruped robot in real-world environments. The results demonstrate that SemGeoNav outperforms existing representative methods, including ViNT and NoMaD, achieving higher success rates and shorter navigation times.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。