用网络视频训练智能体,实现城市级自主导航。
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
- 从海量真实城市视频中提取动作标签,实现无标注模仿学习。
- 在复杂城市环境中导航成功率显著高于现有方法。
- 适合研究城市机器人、自动驾驶与具身智能的开发者。
在动态城市环境中导航对具身智能体构成重大挑战,需具备高级空间推理能力并遵守常识规范。尽管已有进展,现有视觉导航方法在无地图或非道路场景中仍表现不佳,限制了自主代理(如末端配送机器人)的应用。为此,我们提出一种可扩展的数据驱动方法,通过训练数千小时来自网络的真实城市步行与驾驶视频,实现类人城市导航。我们设计了一种简单且可扩展的数据处理流程,从视频中自动提取动作监督信号,实现大规模模仿学习而无需昂贵的人工标注。模型学习到复杂的导航策略,能应对多样化的挑战和关键场景。实验表明,基于大规模、多样化数据集训练显著提升导航性能,优于当前主流方法。本工作展示了利用丰富在线视频数据开发动态城市环境中鲁棒导航策略的潜力。项目主页:https://ai4ce.github.io/CityWalker/。
原文摘要 · Abstract (English)
Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods struggle in map-free or off-street settings, limiting the deployment of autonomous agents like last-mile delivery robots. To overcome these obstacles, we propose a scalable, data-driven approach for human-like urban navigation by training agents on thousands of hours of in-the-wild city walking and driving videos sourced from the web. We introduce a simple and scalable data processing pipeline that extracts action supervision from these videos, enabling large-scale imitation learning without costly annotations. Our model learns sophisticated navigation policies to handle diverse challenges and critical scenarios. Experimental results show that training on large-scale, diverse datasets significantly enhances navigation performance, surpassing current methods. This work shows the potential of using abundant online video data to develop robust navigation policies for embodied agents in dynamic urban settings. Project homepage is at https://ai4ce.github.io/CityWalker/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。