用文本生成全景环境,解决视觉语言导航数据少的问题
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
- 结合扩散模型与领域微调,低成本生成多样全景场景
- 在R2R、R4R、CVDN上分别提升成功率与目标进展
- 适合需要丰富训练环境的VLN研究者使用
视觉-语言导航(VLN)任务要求智能体根据自然语言指令在三维环境中导航,具有广泛应用潜力。然而训练数据稀缺严重制约了该领域进展。本文提出PanoGen++框架,通过生成多样化且相关的全景环境来缓解这一问题。该框架融合预训练扩散模型与领域特定微调,采用低秩适应等高效参数技术以降低计算开销。我们探索两种环境生成方式:掩码图像修复与递归图像外推。前者基于文本描述修复遮蔽区域以最大化新环境生成;后者有助于智能体学习全景中的空间关系。在room-to-room(R2R)、room-for-room(R4R)及协作视觉对话导航(CVDN)数据集上的实证评估显示显著性能提升:R2R测试榜成功率提升2.44%,R4R验证未见集提升0.63%,CVDN验证未见集目标进展增加0.75米。PanoGen++有效增强训练环境的多样性与相关性,从而提升VLN任务的泛化能力与执行效率。
原文摘要 · Abstract (English)
Vision-and-language navigation (VLN) tasks require agents to navigate three-dimensional environments guided by natural language instructions, offering substantial potential for diverse applications. However, the scarcity of training data impedes progress in this field. This paper introduces PanoGen++, a novel framework that addresses this limitation by generating varied and pertinent panoramic environments for VLN tasks. PanoGen++ incorporates pre-trained diffusion models with domain-specific fine-tuning, employing parameter-efficient techniques such as low-rank adaptation to minimize computational costs. We investigate two settings for environment generation: masked image inpainting and recursive image outpainting. The former maximizes novel environment creation by inpainting masked regions based on textual descriptions, while the latter facilitates agents' learning of spatial relationships within panoramas. Empirical evaluations on room-to-room (R2R), room-for-room (R4R), and cooperative vision-and-dialog navigation (CVDN) datasets reveal significant performance enhancements: a 2.44% increase in success rate on the R2R test leaderboard, a 0.63% improvement on the R4R validation unseen set, and a 0.75-meter enhancement in goal progress on the CVDN validation unseen set. PanoGen++ augments the diversity and relevance of training environments, resulting in improved generalization and efficacy in VLN tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。