用生成视频和逆动力学实现无需机器人数据的开放世界导航
ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics
- 分步规划:先分解指令,再生成未来视频轨迹
- 零样本迁移:在真实机器人上直接生效,无需演示数据
- 适合研究通用机器人、视觉语言导航的开发者
让机器人通过自然语言在开放世界中导航是实现通用自主的关键。现有视觉-语言导航依赖昂贵的实体机器人数据训练端到端策略。尽管基于大规模模拟数据的预训练基础模型展现出潜力,但因模拟场景多样性与视觉保真度有限,仍难以规模化和泛化。为此,我们提出ImagiNav,一种新型模块化框架,将视觉规划与机器人执行解耦,可直接利用真实的户外导航视频。该框架分层运行:视觉-语言模型首先将指令分解为文本子目标;微调后的生成视频模型据此想象未来视频轨迹;逆动力学模型从中提取动作序列,由低层控制器跟踪。我们还构建了自动标注的户外导航视频数据流水线,基于逆动力学与预训练视觉-语言模型。ImagiNav在未见机器人上实现强零样本迁移,无需机器人演示,为直接从无标签开放世界数据学习导航的通用机器人铺平道路。
原文摘要 · Abstract (English)
Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While recent foundation models trained on vast simulation data show promise, the challenge of scaling and generalizing due to the limited scene diversity and visual fidelity in simulation persists. To address this gap, we propose ImagiNav, a novel modular paradigm that decouples visual planning from robot actuation, enabling the direct utilization of diverse in-the-wild navigation videos. Our framework operates as a hierarchy: a Vision-Language Model first decomposes instructions into textual subgoals; a finetuned generative video model then imagines the future video trajectory towards that subgoal; finally, an inverse dynamics model extracts the trajectory from the imagined video, which can then be tracked via a low-level controller. We additionally develop a scalable data pipeline of in-the-wild navigation videos auto-labeled via inverse dynamics and a pretrained Vision-Language Model. ImagiNav demonstrates strong zero-shot transfer to robot navigation without requiring robot demonstrations, paving the way for generalist robots that learn navigation directly from unlabeled, open-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。