arXiv:2601.04699cs.ROcs.AI2026-01AAAI被引 5

提出分层规划模型SeqWalker,提升长指令多任务导航能力

SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning

  • 分层规划:高层动态拆解指令,降低认知负荷
  • 低层探索验证策略,自动修正路径错误
  • 在扩展IVLN数据集上验证,性能显著优于现有方法

序列时域视觉-语言导航(SH-VLN)要求智能体根据复杂、长时程的自然语言指令执行一系列多任务导航。当前视觉-语言导航模型在处理此类指令时性能大幅下降,因信息过载削弱了对观测细节的关注能力。为此,我们提出SeqWalker,一种基于分层规划框架的导航模型。其包含:i)高层规划器,根据当前视觉观测动态将全局指令分解为上下文相关的子指令,从而减轻认知负担;ii)低层规划器,采用探索-验证策略,利用指令固有的逻辑结构实现轨迹误差纠正。为评估SH-VLN性能,我们还扩展了IVLN数据集并建立新基准。大量实验表明,所提SeqWalker显著优于现有方法。

原文摘要 · Abstract (English)

Sequential-Horizon Vision-and-Language Navigation (SH-VLN) presents a challenging scenario where agents should sequentially execute multi-task navigation guided by complex, long-horizon language instructions. Current vision-and-language navigation models exhibit significant performance degradation with such multi-task instructions, as information overload impairs the agent's ability to attend to observationally relevant details. To address this problem, we propose SeqWalker, a navigation model built on a hierarchical planning framework. Our SeqWalker features: i) A High-Level Planner that dynamically selects global instructions into contextually relevant sub-instructions based on the agent's current visual observations, thus reducing cognitive load; ii) A Low-Level Planner incorporating an Exploration-Verification strategy that leverages the inherent logical structure of instructions for trajectory error correction. To evaluate SH-VLN performance, we also extend the IVLN dataset and establish a new benchmark. Extensive experiments are performed to demonstrate the superiority of the proposed SeqWalker.

视觉导航语言指令分层规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。