用强化学习让导航模型从看视频变成会体验,提升真实城市中的互动与安全能力。
From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
- 结合离线视频预训练与仿真环境强化学习,兼顾泛化与交互能力。
- 在真实场景重建数据集上实现91.3%的安全通过率,优于基线模型。
- 适合做智能机器人、自动驾驶等需要动态避障的交互式导航任务。
在大规模网络数据上训练的导航基础模型虽能跨环境泛化,但仅依赖离线数据,缺乏对行为后果的推理与反事实理解,难以应对真实城市导航中避障、避人等动态交互需求。为此,我们提出Seeing-to-Experiencing(S2E)学习框架,通过强化学习扩展导航模型能力。S2E融合离线视频预训练与仿真环境后训练,既保留来自大规模真实视频的泛化能力,又通过强化学习增强交互性。关键创新包括:(1) 基于锚点的分布匹配策略,稳定学习并建模多样化运动模式;(2) 残差注意力模块,使模型在仿真中获得反应式行为而不丢失预训练知识。此外,我们构建了基于真实世界场景光场重建的端到端评估基准NavBench-GS,可系统评估模型的泛化性与安全性。
原文摘要 · Abstract (English)
Navigation foundation models trained on massive web-scale data enable agents to generalize across diverse environments and embodiments. However, these models, which are trained solely on offline data, often lack the capacity to reason about the consequences of their actions or adapt through counterfactual understanding. They thus face significant limitations in real-world urban navigation, where interactive and safe behaviors, such as avoiding obstacles and moving pedestrians, are critical. To tackle these challenges, we introduce the Seeing-to-Experiencing (S2E) learning framework to scale the capability of navigation foundation models with reinforcement learning. S2E combines the strengths of pretraining on offline videos and post-training through reinforcement learning. It maintains the model's generalizability acquired from large-scale real-world videos while enhancing its interactivity through reinforcement learning in simulation environments. Specifically, we introduce two innovations: (1) an Anchor-Guided Distribution Matching strategy for offline pretraining, which stabilizes learning and models diverse motion patterns through anchor-based supervision; and (2) a Residual-Attention Module for reinforcement learning, which obtains reactive behaviors from simulation environments without erasing the model's pretrained knowledge. Moreover, we establish a comprehensive end-to-end evaluation benchmark, NavBench-GS, built on photorealistic 3D Gaussian Splatting reconstructions of real-world scenes that incorporate physical interactions. It can systematically assess the generalizability and safety of navigation foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。