用几何规划生成数据,让机器人少靠人教也能学会导航。
Less Is More: Scalable Visual Navigation from Limited Data
- 用经典规划算法生成合成轨迹,补足稀缺真人示范数据。
- 仅用少量专家数据,模型在真实机器人上实现高效导航。
- 适合想降低数据依赖、提升导航泛化能力的研究者。
模仿学习为移动机器人提供了强大的目标导向视觉导航框架,可在避障的同时尊重人类偏好与社交规范。然而其效果高度依赖训练数据的质量与多样性。本文提出利用经典几何规划器生成合成轨迹,以补充昂贵的人类示范数据。我们训练了基于Transformer的视觉导航策略LiMo,从单张RGB图像预测目标条件下的SE(2)轨迹,并发现用规划生成的监督信号增强有限专家示范可显著提升性能。通过消融实验及定性定量分析,我们揭示了数据集规模与多样性对规划性能的影响。实验证明了真实机器人的部署可行性,表明鲁棒视觉导航并非依赖简单增加示范数量,而是通过战略性地构建多样且高质量的数据集实现。结果表明,可扩展的、面向具体实体的几何监督是实现数据高效视觉导航的可行路径。
原文摘要 · Abstract (English)
Imitation learning provides a powerful framework for goal-conditioned visual navigation in mobile robots, enabling obstacle avoidance while respecting human preferences and social norms. However, its effectiveness depends critically on the quality and diversity of training data. In this work, we show how classical geometric planners can be leveraged to generate synthetic trajectories that complement costly human demonstrations. We train Less is More (LiMo), a transformer-based visual navigation policy that predicts goal-conditioned SE(2) trajectories from a single RGB observation, and find that augmenting limited expert demonstrations with planner-generated supervision yields substantial performance gains. Through ablations and complementary qualitative and quantitative analyses, we characterize how dataset scale and diversity affect planning performance. We demonstrate real-robot deployment and argue that robust visual navigation is enabled not by simply collecting more demonstrations, but by strategically curating diverse, high-quality datasets. Our results suggest that scalable, embodiment-specific geometric supervision is a practical path toward data-efficient visual navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。