用模型重标注海量低质数据,训练出能跨环境长距离导航的机器人
Learning to Drive Anywhere with Model-Based Reannotation
- 用短时模型自动为非标注数据生成高质量动作标签
- 在6城3洲实测中实现超300米无地图导航,穿行人群仍稳定
- 适合需要低成本泛化导航能力的机器人研发团队
开发具备广泛泛化能力的机器人视觉导航策略面临重大挑战,主要受限于大规模、多样化训练数据的获取。尽管研究者构建的精选数据集质量高,但规模有限,制约了策略的泛化能力。为此,我们探索利用大量被动采集的数据源,包括海量众包遥控数据和未标注的YouTube视频,尽管其质量可能较低或缺乏动作标签。我们提出基于模型的重标注(MBRA)框架,利用一个学习得到的短时、模型化专家模型,为这些被动数据重新标注或生成高质量动作。这些重标注数据随后被提炼为LogoNav——一种基于视觉目标或GPS航点的长时导航策略。实验表明,使用MBRA处理数据训练的LogoNav达到当前最优性能,在此前未见过的室内外环境中实现超过300米的鲁棒导航。我们在六个城市、三个大洲的机器人车队(包括四足机器人)上进行了广泛的实地评估,验证了该策略在人群密集场景下依然具备良好泛化与导航能力。
原文摘要 · Abstract (English)
Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data. While curated datasets collected by researchers offer high quality, their limited size restricts policy generalization. To overcome this, we explore leveraging abundant, passively collected data sources, including large volumes of crowd-sourced teleoperation data and unlabeled YouTube videos, despite their potential for lower quality or missing action labels. We propose Model-Based ReAnnotation (MBRA), a framework that utilizes a learned short-horizon, model-based expert model to relabel or generate high-quality actions for these passive datasets. This relabeled data is then distilled into LogoNav, a long-horizon navigation policy conditioned on visual goals or GPS waypoints. We demonstrate that LogoNav, trained using MBRA-processed data, achieves state-of-the-art performance, enabling robust navigation over distances exceeding 300 meters in previously unseen indoor and outdoor environments. Our extensive real-world evaluations, conducted across a fleet of robots (including quadrupeds) in six cities on three continents, validate the policy's ability to generalize and navigate effectively even amidst pedestrians in crowded settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。