用双热图预测可行走区域和朝向,让机器人更安全地执行语义导航指令。
Beyond Waypoints: Dual-Heatmap Grounding for Cross-Embodiment Semantic Navigation

- 提出双热图表示法,同时预测可行走区域和朝向约束。
- 在多种机器人上实现92.3%的可达率,显著优于点回归方法。
- 适合需要跨机器人平台安全导航的应用场景。
将开放式的语义指令转化为物理可执行的局部目标是人机交互中的核心挑战。现有导航框架通常回归确定性路径点,这种刚性设定忽略了空间不确定性,常导致目标位于不可通行物体中心,引发严重执行失败。本文聚焦于视野内语义导航场景,机器人接收简短、交错的多模态(文本与图像)提示。为弥合抽象语义意图与物理可达性之间的差距,我们提出一种统一的视觉-语言框架,摒弃单点回归,采用双热图表示:一个导航可达热图捕捉连续可行走区域,另一个朝向热图施加方向约束。这些密集输出本质上构成可微分的语义势场,能无缝集成至下游局部规划器。为此,我们构建了全自动、基于基础模型的合成数据生成管道,并建立全面的仿真基准。大量实验表明,该框架在8B类基线中达到最先进性能。关键特征融合研究及在多种机器人形态(Jetbot, H1, Aliengo)上的仿真分析显示,显式热图预测使可达率(AR)大幅提升。通过将目标可靠放置于可执行自由空间,本框架有效缓解了点回归的脆弱性,为安全的跨体感语义导航提供了可迁移路径。
原文摘要 · Abstract (English)
Grounding open-ended semantic instructions into physically executable local goals is a fundamental challenge in human-robot interaction. While existing navigation frameworks often regress deterministic waypoints, this rigid formulation collapses spatial uncertainty and frequently targets non-traversable object centers, leading to severe execution failures. In this work, we focus on the practical setting of in-FOV semantic navigation, where a robot receives concise, interleaved multimodal (text and image) prompts. To bridge the gap between abstract semantic intent and physical reachability, we propose a unified Vision-Language framework that abandons single-point regression in favor of a Dual-Heatmap representation. Our framework predicts a navigation affordance heatmap that captures continuous reachable regions, coupled with a facing heatmap for orientation constraints. These dense outputs inherently function as a differentiable semantic potential field, integrating seamlessly with downstream local planners. To support this paradigm, we build a fully automated, foundation-model-assisted synthetic data pipeline and establish a comprehensive simulation benchmark. Extensive experiments demonstrate that our framework achieves state-of-the-art performance among comparable 8B baselines. Crucially, a feature-fusion study and simulation studies across diverse robot embodiments (Jetbot, H1, Aliengo) reveal that explicit heatmap prediction drastically improves the Affordance Rate (AR). By placing targets reliably in executable free space, our framework effectively mitigates the brittleness of point regression, offering a transferable path toward safe cross-embodiment semantic navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。