用图像编辑生成关键帧,让机器人更高效地规划动作。
SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution

- 通过连续图像编辑生成任务相关的关键视觉状态
- 在真实场景和未见场景中均提升关键帧预测准确率
- 适合需要高效视觉规划的机器人操作任务
视觉预测已成为具身控制的有前景范式,通过生成未来观测并转化为动作。然而,密集视频生成计算开销大且对多数操作任务不必要,因任务进展可由少数任务相关视觉状态总结。本文研究图像编辑模型能否作为稀疏视觉世界模型用于机器人操作,预测任务级未来状态而无需密集视频推演。在相同机器人数据设置下,对比视频生成模型Wan2.2与图像编辑模型FLUX-Kontext,发现图像编辑生成的任务级关键帧更具可靠性、视觉保真度更高且推理成本显著降低。基于此,提出SWEET——一种单次输入的稀疏视觉规划框架,通过连续图像编辑生成一系列任务相关操作关键帧,条件为语言指令及可选箭头空间引导。随后,目标条件扩散动作预测器将相邻想象关键帧转化为可执行的动作块。为减少真实与编辑视觉子目标间的差异,引入带过滤编辑目标的混合训练策略。在DROID和RoboMimic数据集上的实验表明,SWEET在已见与未见场景中均提升关键帧预测性能,并实现从序列关键帧规划到可执行机器人动作的完整流程,表明图像编辑是具身视觉预测中一个有前景且被低估的方向。
原文摘要 · Abstract (English)
Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. However, dense video generation is computationally expensive and often unnecessary for many manipulation tasks, whose progress can be summarized by a small number of task-relevant visual states. In this work, we study whether image editing models can serve as sparse visual world models for robot manipulation by predicting task-level future states without dense video rollout. We first conduct a controlled comparison between the video generation model Wan2.2 and the image editing model FLUX-Kontext under the same robotic data setting, and find that image editing produces more reliable task-level keyframes with better visual fidelity and substantially lower inference cost. Motivated by this observation, we propose SWEET, a one-shot sparse visual planning framework that progressively generates a sequence of task-relevant manipulation keyframes through successive image editing, conditioned on language instructions and optional arrow-based spatial guidance. A goal-conditioned diffusion action predictor then converts adjacent imagined keyframes into executable action chunks. To reduce the mismatch between real and edited visual subgoals, we further introduce a mixed-training strategy with filtered edited targets. Experiments on DROID and RoboMimic show that SWEET improves keyframe prediction across seen and unseen scenes and enables a full pipeline from sequential keyframe planning to executable robot actions, suggesting that image editing is a promising and underexplored direction for embodied visual prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。