arXiv:2412.09082cs.CV2024-12CVPR被引 91

提出长时序视觉语言导航新任务与评测体系,解决复杂环境下的多阶段规划难题。

Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

  • 设计双向多粒度数据生成平台NavGen,自动生成复杂任务数据
  • 构建含3260个任务、平均150步的LHPR-VLN基准,支持长程评估
  • 提出新指标与动态记忆模块,提升模型在动态环境中的适应性

现有视觉语言导航(VLN)方法多聚焦单阶段任务,难以应对复杂动态环境中的多阶段长时序导航。为此,我们提出长时序视觉语言导航(LH-VLN)新任务,强调跨子任务的长期规划与决策一致性。为支持该任务,我们开发自动化数据生成平台NavGen,采用双向多粒度生成策略构建高质量数据集。同时,构建首个专为长时序导航设计的基准LHPR-VLN,包含3,260个任务,平均每个任务达150步。我们提出独立成功率(ISR)、条件成功率(CSR)及基于真值加权的CSR(CGT)等细粒度评估指标。为增强模型适应性,提出多粒度动态记忆(MGDM)模块,融合短期记忆模糊与长期记忆检索,实现动态环境下的灵活导航。本工作提供完整的数据生成、评测与模型框架,为推动长时序视觉语言导航奠定基础。

原文摘要 · Abstract (English)

Existing Vision-Language Navigation (VLN) methods primarily focus on single-stage navigation, limiting their effectiveness in multi-stage and long-horizon tasks within complex and dynamic environments. To address these limitations, we propose a novel VLN task, named Long-Horizon Vision-Language Navigation (LH-VLN), which emphasizes long-term planning and decision consistency across consecutive subtasks. Furthermore, to support LH-VLN, we develop an automated data generation platform NavGen, which constructs datasets with complex task structures and improves data utility through a bidirectional, multi-granularity generation approach. To accurately evaluate complex tasks, we construct the Long-Horizon Planning and Reasoning in VLN (LHPR-VLN) benchmark consisting of 3,260 tasks with an average of 150 task steps, serving as the first dataset specifically designed for the long-horizon vision-language navigation task. Furthermore, we propose Independent Success Rate (ISR), Conditional Success Rate (CSR), and CSR weight by Ground Truth (CGT) metrics, to provide fine-grained assessments of task completion. To improve model adaptability in complex tasks, we propose a novel Multi-Granularity Dynamic Memory (MGDM) module that integrates short-term memory blurring with long-term memory retrieval to enable flexible navigation in dynamic environments. Our platform, benchmark and method supply LH-VLN with a robust data generation pipeline, comprehensive model evaluation dataset, reasonable metrics, and a novel VLN model, establishing a foundational framework for advancing LH-VLN.

视觉导航长时序多阶段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。