让机器人导航时能主动找新视角纠错,提升成功率。
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
- 基于记忆和错误感知,选择最合适的未探索视角作为纠正路径
- 在两个不同复杂度的数据集上显著提升导航成功率
- 适合需要智能纠错的视觉语言导航场景
具身导航要求机器人根据任务理解并交互环境。视觉语言导航(VLN)是其中一项任务,机器人需根据语言指令和视觉输入,在已见或未见环境中导航。VLN智能体需同时具备局部与全局动作空间:前者用于即时决策,后者用于纠正导航错误。现有方法仅依赖指令-视点对齐进行决策,若不匹配则回溯至先前访问过的视点,但此法易出错,因指令复杂且环境部分可观测。我们提出,回溯次优,具备错误意识的智能体可更高效恢复。最优恢复应扩展至未探索视点(或前沿)。最优前沿是最近观察到但尚未探索、与指令一致且具有新颖性的视点。本文提出名为StratXplore的记忆型、错误感知路径规划策略,实现全局与局部动作规划,以选择最佳前沿进行路径修正。该方法在导航过程中收集所有历史动作与视点特征,并据此选出适于恢复的最优前沿。实验结果表明,这一简单而有效的方法在两个不同任务复杂度的VLN数据集上均提升了成功率达显著水平。
原文摘要 · Abstract (English)
Embodied navigation requires robots to understand and interact with the environment based on given tasks. Vision-Language Navigation (VLN) is an embodied navigation task, where a robot navigates within a previously seen and unseen environment, based on linguistic instruction and visual inputs. VLN agents need access to both local and global action spaces; former for immediate decision making and the latter for recovering from navigational mistakes. Prior VLN agents rely only on instruction-viewpoint alignment for local and global decision making and back-track to a previously visited viewpoint, if the instruction and its current viewpoint mismatches. These methods are prone to mistakes, due to the complexity of the instruction and partial observability of the environment. We posit that, back-tracking is sub-optimal and agent that is aware of its mistakes can recover efficiently. For optimal recovery, exploration should be extended to unexplored viewpoints (or frontiers). The optimal frontier is a recently observed but unexplored viewpoint that aligns with the instruction and is novel. We introduce a memory-based and mistake-aware path planning strategy for VLN agents, called \textit{StratXplore}, that presents global and local action planning to select the optimal frontier for path correction. The proposed method collects all past actions and viewpoint features during navigation and then selects the optimal frontier suitable for recovery. Experimental results show this simple yet effective strategy improves the success rate on two VLN datasets with different task complexities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。