arXiv:2511.00033cs.ROcs.AI2025-11NeurIPS被引 4

让智能体按空间结构和任务目标导航,零样本效果显著提升

STRIDER: Navigation via Instruction-Aligned Structural Decision Space Optimization

  • 用空间结构约束动作选择,动态调整行为以匹配任务进度
  • 在R2R-CE和RxR-CE上成功率从29%提升至35%,相对增益20.7%
  • 适合做零样本视觉语言导航的算法研究者与工程落地应用

零样本视觉-语言导航在连续环境(VLN-CE)中要求智能体在未见过的3D场景中,仅凭自然语言指令完成导航,且不进行场景特定训练。该任务的核心挑战在于:长时程执行中,如何保证智能体的行为既符合空间结构,又契合任务意图。现有方法因缺乏结构化决策机制和对历史动作反馈的不足整合,常表现不佳。为此,我们提出STRIDER(指令对齐的结构化决策空间优化框架),通过融合空间布局先验与动态任务反馈,系统优化智能体的决策空间。本方法包含两项关键创新:1)结构化航点生成器,利用空间结构约束动作空间;2)任务对齐调节器,基于任务进展动态调整行为,确保全程语义一致性。在R2R-CE和RxR-CE基准上的大量实验表明,STRIDER显著超越当前最优方法,在关键指标上实现显著提升:成功率(SR)由29%提升至35%,相对增益达20.7%。结果凸显了空间约束决策与反馈引导执行在提升零样本VLN-CE导航保真度中的重要性。

原文摘要 · Abstract (English)

The Zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) task requires agents to navigate previously unseen 3D environments using natural language instructions, without any scene-specific training. A critical challenge in this setting lies in ensuring agents' actions align with both spatial structure and task intent over long-horizon execution. Existing methods often fail to achieve robust navigation due to a lack of structured decision-making and insufficient integration of feedback from previous actions. To address these challenges, we propose STRIDER (Instruction-Aligned Structural Decision Space Optimization), a novel framework that systematically optimizes the agent's decision space by integrating spatial layout priors and dynamic task feedback. Our approach introduces two key innovations: 1) a Structured Waypoint Generator that constrains the action space through spatial structure, and 2) a Task-Alignment Regulator that adjusts behavior based on task progress, ensuring semantic alignment throughout navigation. Extensive experiments on the R2R-CE and RxR-CE benchmarks demonstrate that STRIDER significantly outperforms strong SOTA across key metrics; in particular, it improves Success Rate (SR) from 29% to 35%, a relative gain of 20.7%. Such results highlight the importance of spatially constrained decision-making and feedback-guided execution in improving navigation fidelity for zero-shot VLN-CE.

视觉导航语言导航强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。