大规模数据提升导航泛化能力,多样性比数量更重要。
Data Scaling for Navigation in Unknown Environments
- 用4565小时跨161地数据训练端到端视觉导航模型。
- 数据多样性提升15%导航准确率,数量增益很快饱和。
- 噪声数据下回归模型优于复杂生成架构,适合实操部署。
模仿学习的导航策略在未见环境中泛化仍是重大挑战。我们首次系统研究了数据量与多样性对端到端、无地图视觉导航真实世界泛化的影响。基于覆盖35个国家161个地点的4,565小时众包数据集,训练点目标导航策略,并在四个国家的路边机器人上评估闭环控制性能,覆盖125公里自主驾驶。结果表明,大规模训练数据可实现未知环境下的零样本导航,接近专有环境训练策略的性能。关键发现:数据多样性远比数据量重要——训练集地理位置翻倍使导航误差降低约15%,而增加已有地点数据的收益迅速饱和。此外,在存在噪声的众包数据下,简单的回归模型表现优于生成和序列架构。代码、评估设置及示例视频已公开。
原文摘要 · Abstract (English)
Generalization of imitation-learned navigation policies to environments unseen in training remains a major challenge. We address this by conducting the first large-scale study of how data quantity and data diversity affect real-world generalization in end-to-end, map-free visual navigation. Using a curated 4,565-hour crowd-sourced dataset collected across 161 locations in 35 countries, we train policies for point goal navigation and evaluate their closed-loop control performance on sidewalk robots operating in four countries, covering 125 km of autonomous driving. Our results show that large-scale training data enables zero-shot navigation in unknown environments, approaching the performance of policies trained with environment-specific demonstrations. Critically, we find that data diversity is far more important than data quantity. Doubling the number of geographical locations in a training set decreases navigation errors by ~15%, while performance benefit from adding data from existing locations saturates with very little data. We also observe that, with noisy crowd-sourced data, simple regression-based models outperform generative and sequence-based architectures. We release our policies, evaluation setup and example videos at https://lasuomela.github.io/navigation_scaling/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。