X-Mobility让机器人在复杂环境中端到端自适应导航,无需重训即可跨场景、跨设备通用。
X-MOBILITY: End-To-End Generalizable Navigation via World Modeling
- 用自回归世界建模+隐状态空间捕捉环境动态变化
- 多头解码器学习强关联导航能力的丰富状态表示
- 分离环境建模与动作策略,支持无监督与有监督数据联合训练
复杂环境下通用导航仍是机器人领域的重大挑战。传统方法在杂乱场景中表现差且需大量调参,学习类方法难以泛化至分布外环境。本文提出X-Mobility,一种端到端可泛化的导航模型,通过三个关键设计克服现有局限:首先,采用自回归世界建模架构与隐状态空间以捕捉环境动态;其次,多样化的多头解码器使模型学习到与高效导航技能强相关的丰富状态表征;第三,将世界建模与动作策略解耦,可有效利用多种数据源——离策略数据用于学习环境动态,同策略带监督数据则用于优化动作策略。大量实验表明,X-Mobility不仅具备优异泛化能力,还超越当前最先进导航方法。此外,其实现零样本仿真到现实(Sim2Real)迁移,并展现出强大的跨机体泛化潜力。
原文摘要 · Abstract (English)
General-purpose navigation in challenging environments remains a significant problem in robotics, with current state-of-the-art approaches facing myriad limitations. Classical approaches struggle with cluttered settings and require extensive tuning, while learning-based methods face difficulties generalizing to out-of-distribution environments. This paper introduces X-Mobility, an end-to-end generalizable navigation model that overcomes existing challenges by leveraging three key ideas. First, X-Mobility employs an auto-regressive world modeling architecture with a latent state space to capture world dynamics. Second, a diverse set of multi-head decoders enables the model to learn a rich state representation that correlates strongly with effective navigation skills. Third, by decoupling world modeling from action policy, our architecture can train effectively on a variety of data sources, both with and without expert policies: off-policy data allows the model to learn world dynamics, while on-policy data with supervisory control enables optimal action policy learning. Through extensive experiments, we demonstrate that X-Mobility not only generalizes effectively but also surpasses current state-of-the-art navigation approaches. Additionally, X-Mobility also achieves zero-shot Sim2Real transferability and shows strong potential for cross-embodiment generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。