arXiv:2512.09607cs.ROcs.CV2025-12中稿 · AAAI被引 8

用网络视频训练机器人在真实城市跟自然语言导航

UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories

  • 基于海量城市步行视频,对齐语言指令与真实地标轨迹
  • 构建超1500小时数据,300万条指令-轨迹-地标三元组
  • 可处理噪声指令,泛化至未见城市环境,适合配送机器人

使用自然语言指令导航复杂城市环境对具身智能体构成重大挑战,包括语言指令噪声、空间指代模糊、地标多样性和动态街景。现有视觉导航方法通常局限于模拟或非街道环境,且依赖精确目标格式(如坐标或图像),限制了其在陌生城市中执行末段配送任务的自主机器人应用。为此,我们提出UrbanNav,一个可扩展的框架,使具身智能体能在多样化城市环境中遵循自由格式语言指令。利用网络规模的城市步行视频,我们开发了一种可扩展的标注流程,将人类导航轨迹与基于真实地标语言指令对齐。UrbanNav包含超过1500小时的导航数据和300万条指令-轨迹-地标三元组,涵盖广泛的城市场景。模型学习到鲁棒的导航策略,展现出卓越的空间推理能力、对噪声指令的鲁棒性,并能泛化至未见城市环境。实验结果表明,UrbanNav显著优于现有方法,凸显大规模网络视频数据在实现语言引导的真实世界城市导航方面的潜力。

原文摘要 · Abstract (English)

Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dynamic street scenes. Current visual navigation methods are typically limited to simulated or off-street environments, and often rely on precise goal formats, such as specific coordinates or images. This limits their effectiveness for autonomous agents like last-mile delivery robots navigating unfamiliar cities. To address these limitations, we introduce UrbanNav, a scalable framework that trains embodied agents to follow free-form language instructions in diverse urban settings. Leveraging web-scale city walking videos, we develop an scalable annotation pipeline that aligns human navigation trajectories with language instructions grounded in real-world landmarks. UrbanNav encompasses over 1,500 hours of navigation data and 3 million instruction-trajectory-landmark triplets, capturing a wide range of urban scenarios. Our model learns robust navigation policies to tackle complex urban scenarios, demonstrating superior spatial reasoning, robustness to noisy instructions, and generalization to unseen urban settings. Experimental results show that UrbanNav significantly outperforms existing methods, highlighting the potential of large-scale web video data to enable language-guided, real-world urban navigation for embodied agents.

城市导航语言引导具身智能多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。