让机器人在城市中可靠导航,能听懂长路线指令并避开障碍。
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- 用视觉-语言-动作模型对齐路线点与真实画面,动态规划路径。
- 在MetaUrban上比基线高出55%以上,实测可稳定运行于真实城市环境。
- 适合做配送机器人、自动驾驶等需要长程复杂导航的场景。
城市微出行应用(如配送机器人)要求在大规模城市环境中可靠导航,并遵循长程路线指令。这一任务极具挑战性,因真实城市环境动态且无序,而现有导航方法多局限于短距离可控场景。有效城市微出行需兼顾低层能力(如到达目标点、避障)与高层能力(如路线与视觉对齐)。为此,我们提出UrbanVLA,一种面向可扩展城市导航的路线条件化视觉-语言-动作(VLA)框架。该方法在执行过程中显式对齐噪声路线点与视觉观测,并据此规划轨迹。为使UrbanVLA掌握双层导航技能,我们采用两阶段训练流程:首先在仿真环境和从网络视频解析的轨迹上进行监督微调(SFT);随后在仿真与真实数据混合集上进行强化微调(RFT),提升模型在真实场景中的安全性和适应性。实验表明,UrbanVLA在MetaUrban的SocialNav任务上超越强基线超过55%。此外,其在真实世界中实现了可靠导航,展现出对大规模城市环境的可扩展性及对现实不确定性的鲁棒性。
原文摘要 · Abstract (English)
Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This task is particularly challenging due to the dynamic and unstructured nature of real-world city areas, yet most existing navigation methods remain tailored to short-scale and controllable scenarios. Effective urban micromobility requires two complementary levels of navigation skills: low-level capabilities such as point-goal reaching and obstacle avoidance, and high-level capabilities, such as route-visual alignment. To this end, we propose UrbanVLA, a route-conditioned Vision-Language-Action (VLA) framework designed for scalable urban navigation. Our method explicitly aligns noisy route waypoints with visual observations during execution, and subsequently plans trajectories to drive the robot. To enable UrbanVLA to master both levels of navigation, we employ a two-stage training pipeline. The process begins with Supervised Fine-Tuning (SFT) using simulated environments and trajectories parsed from web videos. This is followed by Reinforcement Fine-Tuning (RFT) on a mixture of simulation and real-world data, which enhances the model's safety and adaptability in real-world settings. Experiments demonstrate that UrbanVLA surpasses strong baselines by more than 55% in the SocialNav task on MetaUrban. Furthermore, UrbanVLA achieves reliable real-world navigation, showcasing both scalability to large-scale urban environments and robustness against real-world uncertainties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。