用视觉流生成通用操作规划,仅靠语言和图像就能完成复杂长程任务。
FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model
- 以视觉流为动作表示,通过三模块协同实现基于视觉空间的规划。
- 在多个基准上提升长程视频计划成功率与质量,最高达87%成功率。
- 适合需要长程规划的机器人任务,可指导底层控制策略训练。
我们致力于构建一个可随模型与数据规模扩展的基于模型的规划框架,用于通用操作任务,仅需语言和视觉输入。为此,提出流中心生成规划(FLIP),其包含三个核心模块:1. 多模态流生成模型作为通用动作提议模块;2. 流条件视频生成模型作为动态模块;3. 视觉-语言表征学习模型作为价值模块。给定初始图像与语言指令作为目标,FLIP 可逐步搜索最大化折扣回报的长程流与视频计划以完成任务。FLIP 能在跨物体、机器人与任务间合成长程计划,以图像流作为通用动作表示,密集流信息为长程视频生成提供丰富引导。此外,生成的流与视频计划可指导低层控制策略的训练。在多样基准上的实验表明,FLIP 提升了长程视频计划合成的成功率与质量,并具备交互式世界模型特性,为后续研究开辟更广泛应用。视频演示见官网:https://nus-lins-lab.github.io/flipweb/。
原文摘要 · Abstract (English)
We aim to develop a model-based planning framework for world models that can be scaled with increasing model and data budgets for general-purpose manipulation tasks with only language and vision inputs. To this end, we present FLow-centric generative Planning (FLIP), a model-based planning algorithm on visual space that features three key modules: 1. a multi-modal flow generation model as the general-purpose action proposal module; 2. a flow-conditioned video generation model as the dynamics module; and 3. a vision-language representation learning model as the value module. Given an initial image and language instruction as the goal, FLIP can progressively search for long-horizon flow and video plans that maximize the discounted return to accomplish the task. FLIP is able to synthesize long-horizon plans across objects, robots, and tasks with image flows as the general action representation, and the dense flow information also provides rich guidance for long-horizon video generation. In addition, the synthesized flow and video plans can guide the training of low-level control policies for robot execution. Experiments on diverse benchmarks demonstrate that FLIP can improve both the success rates and quality of long-horizon video plan synthesis and has the interactive world model property, opening up wider applications for future works.Video demos are on our website: https://nus-lins-lab.github.io/flipweb/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。