arXiv:2603.09961cs.ROcs.AI2026-03

在遮挡情况下,根据语言指令预测机器人可通行目标位置。

BEACON: Language-Conditioned Navigation Affordance Prediction under Occlusion

  • 将视觉语言模型与深度信息融合,在鸟瞰图空间生成可通行热力图。
  • 在遮挡目标上准确率比当前最优方法提升22.74个百分点。
  • 适合需要复杂环境导航的移动机器人研究者使用。

语言条件下的局部导航要求机器人从当前观测和开放式词汇、关系性指令中推断出附近的可通行目标位置。现有视觉-语言空间定位方法通常依赖视觉语言模型(VLM)在图像空间推理,生成与可见像素绑定的二维预测,因此难以推断被家具或移动人类遮挡区域的目标位置。为此,我们提出BEACON,通过在包含遮挡区域的局部范围内生成以机器人为中心的鸟瞰图(BEV)可通行热力图来解决该问题。给定指令和机器人四周方向的环绕视图RGB-D观测,BEACON通过向VLM注入空间线索,并融合VLM输出与深度导出的BEV特征,生成热力图。我们在Habitat模拟器构建的遮挡感知数据集上进行了详尽实验,验证了BEV空间建模及各模块设计的有效性。在包含遮挡目标的验证子集上,本方法在几何距离阈值上的平均准确率较当前最优图像空间基线提升22.74个百分点。

原文摘要 · Abstract (English)

Language-conditioned local navigation requires a robot to infer a nearby traversable target location from its current observation and an open-vocabulary, relational instruction. Existing vision-language spatial grounding methods usually rely on vision-language models (VLMs) to reason in image space, producing 2D predictions tied to visible pixels. As a result, they struggle to infer target locations in occluded regions, typically caused by furniture or moving humans. To address this issue, we propose BEACON, which predicts an ego-centric Bird's-Eye View (BEV) affordance heatmap over a bounded local region including occluded areas. Given an instruction and surround-view RGB-D observations from four directions around the robot, BEACON predicts the BEV heatmap by injecting spatial cues into a VLM and fusing the VLM's output with depth-derived BEV features. Using an occlusion-aware dataset built in the Habitat simulator, we conduct detailed experimental analysis to validate both our BEV space formulation and the design choices of each module. Our method improves the accuracy averaged across geodesic thresholds by 22.74 percentage points over the state-of-the-art image-space baseline on the validation subset with occluded target locations. Our project page is: https://xin-yu-gao.github.io/beacon.

导航视觉语言遮挡处理鸟瞰图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。