通过融合视觉与目标图像,提升机器人导航的泛化能力。
PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
- 用预训练ViT融合视觉与目标图像,增强全局导航表征
- 零样本场景下性能提升22.6%,微调场景提升37.5%
- 减少标注数据依赖,适合真实场景部署
近期研究探索了用于视觉导航的预训练(基础)模型,旨在实现跨环境的通用导航与正向迁移,并提升未见场景下的零样本性能。本文提出PIG-Nav(预训练图像目标导航),深入研究视觉导航模型的预训练策略,在两个关键方面做出贡献。模型层面,我们识别出两项关键设计:(1) 采用早期融合网络结构,通过适当预训练的ViT图像编码器结合视觉观测与目标图像;(2) 引入合适辅助任务以增强全局导航表征学习,从而进一步提升导航性能。数据层面,我们提出一种新型数据预处理流程,高效标注大规模游戏视频数据集用于导航模型训练。实验表明,将多样化游戏视频加入现有开放导航数据集可显著提升模型表现。在两个复杂仿真环境和一个真实世界环境中,本模型相较现有视觉导航基础模型,在零样本设置下平均提升22.6%,微调设置下提升37.5%。结果推动了预训练图像目标导航模型的性能边界。值得注意的是,该模型在保持竞争力的同时,显著降低微调数据需求,展现出在真实场景中仅需少量标注监督即可部署的潜力。
原文摘要 · Abstract (English)
Recent studies have explored pretrained (foundation) models for vision-based robotic navigation, aiming to achieve generalizable navigation and positive transfer across diverse environments while enhancing zero-shot performance in unseen settings. In this work, we introduce PIG-Nav (Pretrained Image-Goal Navigation), a new approach that further investigates pretraining strategies for vision-based navigation models and contributes in two key areas. Model-wise, we identify two critical design choices that consistently improve the performance of pretrained navigation models: (1) integrating an early-fusion network structure to combine visual observations and goal images via appropriately pretrained Vision Transformer (ViT) image encoder, and (2) introducing suitable auxiliary tasks to enhance global navigation representation learning, thus further improving navigation performance. Dataset-wise, we propose a novel data preprocessing pipeline for efficiently labeling large-scale game video datasets for navigation model training. We demonstrate that augmenting existing open navigation datasets with diverse gameplay videos improves model performance. Our model achieves an average improvement of 22.6% in zero-shot settings and a 37.5% improvement in fine-tuning settings over existing visual navigation foundation models in two complex simulated environments and one real-world environment. These results advance the state-of-the-art in pretrained image-goal navigation models. Notably, our model maintains competitive performance while requiring significantly less fine-tuning data, highlighting its potential for real-world deployment with minimal labeled supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。