让机器人在陌生城市自主导航,干预次数减少70%
CREStE: Scalable Mapless Navigation with Internet Scale Priors and Counterfactual Guidance
- 用视觉基础模型学习通用环境感知表征
- 通过反事实轨迹推理关键导航线索,减少人为标注依赖
- 长距离无地图导航表现优异,适合真实城市场景
我们提出CREStE,一种可扩展的基于学习的无地图导航框架,旨在解决室外城市导航中的开放世界泛化与鲁棒性挑战。核心在于学习能泛化至开放集因素(如新语义类别、地形、动态物体)的感知表征,并从有限示范中推断专家对齐的导航代价。CREStE引入两项关键技术:1)基于视觉基础模型(VFM)蒸馏的目标,用于学习结构化的鸟瞰视角感知表征;2)反事实逆强化学习(Counterfactual IRL),一种新颖的主动学习范式,利用反事实轨迹示范来推理导航代价推断中最关键的线索。我们在多种城市、非铺装及住宅环境中评估了千米级无地图导航任务,结果表明,CREStE在所有最先进方法中表现最佳,人类干预次数减少70%。其中,在未见过的环境中完成2公里任务仅需1次干预,充分展示其在长时程无地图导航中的鲁棒性与有效性。视频与补充材料见项目页:https://amrl.cs.utexas.edu/creste
原文摘要 · Abstract (English)
We introduce CREStE, a scalable learning-based mapless navigation framework to address the open-world generalization and robustness challenges of outdoor urban navigation. Key to achieving this is learning perceptual representations that generalize to open-set factors (e.g. novel semantic classes, terrains, dynamic entities) and inferring expert-aligned navigation costs from limited demonstrations. CREStE addresses both these issues, introducing 1) a visual foundation model (VFM) distillation objective for learning open-set structured bird's-eye-view perceptual representations, and 2) counterfactual inverse reinforcement learning (IRL), a novel active learning formulation that uses counterfactual trajectory demonstrations to reason about the most important cues when inferring navigation costs. We evaluate CREStE on the task of kilometer-scale mapless navigation in a variety of city, offroad, and residential environments and find that it outperforms all state-of-the-art approaches with 70% fewer human interventions, including a 2-kilometer mission in an unseen environment with just 1 intervention; showcasing its robustness and effectiveness for long-horizon mapless navigation. Videos and additional materials can be found on the project page: https://amrl.cs.utexas.edu/creste
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。