自动生成带细粒度对齐的导航指令,提升智能体导航准确性。
Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations
- 通过分段轨迹与多模型协作生成子指令-轨迹对,实现细粒度对齐。
- 在多个主流导航模型上提升性能,子指令对齐使决策更精准。
- 适合研究视觉语言导航与跨模态对齐的学者,数据可直接复用。
视觉-语言导航(VLN)使智能体通过融合视觉感知与自然语言指令实现环境导航,但受限于细粒度跨模态对齐标注稀缺。现有数据集多聚焦全局指令-轨迹匹配,忽视子指令级与实体级对齐,影响导航动作决策。为此,本文提出FCA-NIG生成框架,自动构建带双层级细粒度对齐标注的导航指令。该框架首先将轨迹分段,经GLIP检测地标、人工构造指令、OFA-Speaker生成类似R2R的指令,并由CLIP选择实体,形成带实体-地标标注的子指令-轨迹对;最终聚合为完整指令-轨迹对。由此生成的FCA-R2R数据集是首个大规模具备精确子指令-子轨迹与实体-地标对齐的增强数据集。实验表明,使用FCA-R2R训练可显著提升多种前沿VLN代理(如SF、EnvDrop、RecBERT、HAMT)性能。子指令-轨迹对齐增强代理状态感知与决策准确率,实体-地标对齐进一步提升导航表现与泛化能力。结果验证了FCA-NIG在无需人工标注下生成高质量、可扩展训练数据的有效性,推动复杂导航任务中的细粒度跨模态学习发展。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) enables intelligent agents to navigate environments by integrating visual perception and natural language instructions, yet faces significant challenges due to the scarcity of fine-grained cross-modal alignment annotations. Existing datasets primarily focus on global instruction-trajectory matching, neglecting sub-instruction-level and entity-level alignments critical for accurate navigation action decision-making. To address this limitation, we propose FCA-NIG, a generative framework that automatically constructs navigation instructions with dual-level fine-grained cross-modal annotations. In this framework, an augmented trajectory is first divided into sub-trajectories, which are then processed through GLIP-based landmark detection, crafted instruction construction, OFA-Speaker based R2R-like instruction generation, and CLIP-powered entity selection, generating sub-instruction-trajectory pairs with entity-landmark annotations. Finally, these sub-pairs are aggregated to form a complete instruction-trajectory pair. The framework generates the FCA-R2R dataset, the first large-scale augmentation dataset featuring precise sub-instruction-sub-trajectory and entity-landmark alignments. Extensive experiments demonstrate that training with FCA-R2R significantly improves the performance of multiple state-of-the-art VLN agents, including SF, EnvDrop, RecBERT, and HAMT. Incorporating sub-instruction-trajectory alignment enhances agents' state awareness and decision accuracy, while entity-landmark alignment further boosts navigation performance and generalization. These results highlight the effectiveness of FCA-NIG in generating high-quality, scalable training data without manual annotation, advancing fine-grained cross-modal learning in complex navigation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。