解决视觉语言导航中最后3米目标定位不准的问题
From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

- 提出可视性感知的停止策略,避免过早终止
- 在4个模型上提升目标接近精度和可见性成功率
- 适合关注细粒度目标定位的导航研究者
视觉-语言导航(VLN)代理在传统评估标准下可能达标,却仍无法建立可靠的物体级定位,因现有评估主要奖励在3米范围内停止,忽略最终朝向和目标可见性。我们定义此问题为‘最后3米定位差距’,并引入三项以实例为中心的指标:接近精度、目标可见性和最终视角定位。为此,我们提出REALM(区域到实体对齐),一种无需修改架构的即插即用模块,将细粒度目标接近与长距离导航解耦。REALM采用可视性感知的停止策略,减少过早终止并改善最终视角对齐。我们还构建了REVERIE-AIM数据集,提供对象实例级目标和18万条短程训练样本用于最后阶段目标接近。在四个不同VLN骨干模型上的广泛评估显示,REALM持续提升接近精度和视觉定位成功率,证明其广泛适用性。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation protocols mainly reward stopping within a 3-meter radius and largely ignore the agent's final orientation and target visibility. We formalize this limitation as the Last-3-Meter Grounding Gap and introduce three instance-centric metrics to quantify proximity precision, target visibility, and final-view grounding. To mitigate this gap, we propose REALM (Region-to-Entity Alignment for Last-3-Meter Navigation), a plug-and-play, architecture-agnostic refinement module that decouples fine-grained target approaching from long-horizon navigation. REALM uses a visibility-aware stopping strategy to reduce premature termination and improve final viewpoint alignment. We further construct REVERIE-AIM, which provides object-instance-level goals and 180K short-horizon training samples for final-stage target approaching. Extensive evaluations across four diverse VLN backbones show that REALM consistently improves proximity precision and visual grounding success, demonstrating its broad applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。