arXiv:2605.31121cs.ROcs.AI2026-05

解决户外导航中语义线索中断导致的迷失问题,让机器人在无指引时仍能稳定前行。

TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues

论文配图:TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues
图 1 · 摘自论文原文
  • 用实时可通行性地图将语义方向转为可行路径,避免无效前进。
  • 在600-1000米路线中,仿真成功率提升超10个百分点,实测成功率达40%。
  • 适合四足和轮式机器人,在长时间无线索环境下表现更稳健。

长距离开放环境中的户外视觉语言导航常因语义线索中断而受扰,即目标提示变得稀疏、被遮挡或移出视野。一旦线索消失,智能体进入无提示阶段,常出现回溯、徘徊或盲目探索。现有基于记忆的方法虽能缓解此问题,但在需绕行不可通行区域时仍会失败:记忆中的线索方向可能不可达,迫使迂回路径,延长无提示期,导致以机器人为中心的线索过时,历史信息模糊。因此,可通行性不仅是局部安全考量,更是维持目标导向引导的稳定性条件。本文提出统一框架,通过保持可通行性一致的可执行引导,使智能体在长时间无提示阶段仍能持续前行。具体地,方法从可见的目标或探索线索中提取语义方位,并结合实时近场可通行性剖面,生成与目标一致且可行的行动方向,超越仅排除障碍的安全过滤。为防止绕行中引导退化,将间断的二维证据提升至与世界对齐的三维线索记忆,采用不确定性感知读取机制,确保引导始终可达且稳定。在四足与轮式平台上的600–1000米路线评估中,该方法使仿真成功率超过最强基线10个百分点,实测成功率达40%,远高于基线的17.5%,且在长时间无线索间隔下展现出显著更强鲁棒性。

原文摘要 · Abstract (English)

Outdoor vision-language navigation (VLN) in long-range, open-world environments is frequently disrupted by semantic-cue interruptions, where informative goal cues become sparse, occluded, or leave the field of view. Once such cues disappear, agents enter a cue-free phase and often degrade into backtracking, oscillatory headings, or aimless exploration. While memory-based methods attempt to bridge these gaps, they often fail under traversability-driven detours: the remembered cue direction may be infeasible, forcing detours that prolong cue-free phases and gradually render robot-centric cues stale and implicit histories blurred. This makes traversability a stability condition for maintaining goal-directed guidance, rather than merely a local safety concern. We propose a unified outdoor VLN framework that survives semantic-cue interruptions by maintaining traversability-consistent executable guidance throughout prolonged cue-free phases. Specifically, our method extracts semantic bearings from visibility-gated goal or exploration cues and grounds them into executable headings using a real-time near-field traversability profile, providing goal-consistent feasible guidance beyond reject-only safety filtering. To prevent guidance degradation during detours, we lift intermittent 2D evidence into a world-aligned 3D cue memory with an uncertainty-aware readout mechanism, ensuring guidance remains continuously reachable and stable as the robot moves. We evaluate the framework on quadrupedal and wheeled platforms over 600--1000 m routes. Our method improves simulation success rate by over 10 percentage points over the strongest baseline and achieves a real-world success rate of 40%, compared to 17.5% for the strongest baseline, with substantially higher robustness during prolonged cue-free intervals.

户外导航视觉语言记忆增强可通行性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。