通过自迭代生成探索轨迹,提升视觉语言导航的泛化能力
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
- 用短路径数据训练初始智能体,再生成新探索轨迹用于迭代优化
- 在SOON未见验证集上达50.9%成功率,比之前方法高13.9个百分点
- 适合需要强探索和跨任务迁移的视觉语言导航研究者
目标导向的视觉-语言导航要求智能体在未知环境中自主探索以到达指定目标,而无需逐步指令。现有方法多依赖最短路径轨迹,缺乏有效的探索先验。为此,我们提出SID——一种基于自改进演示的目标导向视觉-语言导航学习方法。具体而言,SID首先在采样自环境的最短路径数据上训练初始智能体,随后利用该智能体生成新的探索轨迹作为高质量示范。这些新示范提供更强探索策略,用于训练更优智能体,后者又生成更高品质示范进入下一轮训练。该迭代自改进流程可轻松扩展至新环境,生成的示范具有高度可迁移性,显著提升各类视觉-语言导航任务的性能上限。大量实验表明,SID大幅增强导航智能体的探索能力和泛化性,在包括REVERIE、SOON在内的多个基准上达到新最佳表现,并展现出对物体目标导航和VLN-CE任务的强迁移能力。其在未见的SOON验证集上取得50.9%的成功率,超越先前领先方法13.9个百分点。
原文摘要 · Abstract (English)
Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instructions. Existing methods tend to exclusively utilize shortest-path trajectories, lacking effective exploration priors for training navigation agents. To address the above challenges, we present SID, a goal-oriented vision-and-language navigation learning approach with Self-Improving Demonstrations. Specifically, SID learns an initial agent on the shortest-path data sampled from environments and then leverages this agent to generate novel exploration trajectories. The novel rollouts provide demonstrations with stronger exploration strategies to train a better agent, which in turn produces higher-quality agent demonstrations for the next round of training. We show that this iterative self-improving pipeline readily scales to new environments, and the resulting demonstrations are highly transferable, elevating the performance ceiling across a variety of vision-and-language navigation tasks. Extensive experiments demonstrate that SID significantly boosts the exploration capabilities and generalization of navigation agents. The resulting agent achieves new state-of-the-art performance on goal-oriented vision-and-language navigation benchmarks, including REVERIE, SOON as well as strong transferability to object-goal navigation and VLN-CE. It notably achieves a 50.9% success rate on the unseen validation splits of SOON, surpassing prior leading approaches by a margin of 13.9%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。