提出新数据生成方法,让导航智能体在未见环境中表现更稳定。
World-Consistent Data Generation for Vision-and-Language Navigation
- 分两阶段生成:先保证视角间空间连贯性,再确保观察全景一致性。
- 在多个数据集上实现最新最好效果,显著提升未见环境的泛化能力。
- 适合研究视觉语言导航与数据增强的学者,尤其关注真实场景泛化。
视觉-语言导航(VLN)要求智能体根据自然语言指令在逼真的环境中导航。现有主要挑战是数据稀缺,导致对未见环境泛化性能差。尽管数据增强是扩充数据集的可行路径,但如何同时保证数据多样性和世界一致性仍具挑战。为此,我们提出世界一致的数据生成(WCGEN)框架,兼顾多样性与世界一致性,以提升智能体在新环境中的泛化能力。该框架分为两个阶段:轨迹阶段利用基于点云的技术确保视角间的空间一致性;视角阶段采用新颖的角度合成方法,保障整体观察的空间与环绕一致性。通过结合3D知识精准预测视角变化,本方法在生成过程中保持世界一致性。在多种数据集上的实验验证了其有效性,表明该数据增强策略使智能体在所有导航任务中达到新最优性能,并显著增强其对未见环境的泛化能力。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through photorealistic environments following natural-language instructions. One main obstacle existing in VLN is data scarcity, leading to poor generalization performance over unseen environments. Though data argumentation is a promising way for scaling up the dataset, how to generate VLN data both diverse and world-consistent remains problematic. To cope with this issue, we propose the world-consistent data generation (WCGEN), an efficacious data-augmentation framework satisfying both diversity and world-consistency, aimed at enhancing the generalization of agents to novel environments. Roughly, our framework consists of two stages, the trajectory stage which leverages a point-cloud based technique to ensure spatial coherency among viewpoints, and the viewpoint stage which adopts a novel angle synthesis method to guarantee spatial and wraparound consistency within the entire observation. By accurately predicting viewpoint changes with 3D knowledge, our approach maintains the world-consistency during the generation procedure. Experiments on a wide range of datasets verify the effectiveness of our method, demonstrating that our data augmentation strategy enables agents to achieve new state-of-the-art results on all navigation tasks, and is capable of enhancing the VLN agents' generalization ability to unseen environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。