arXiv:2512.00532cs.CVcs.RO2025-12被引 3

用预训练图像模型生成机器人操作视频,无需复杂训练。

Image Generation as a Visual Planner for Robotic Manipulation

  • 用文本或轨迹条件控制图像模型生成连贯操作视频。
  • 在Jaco Play、Bridge V2等数据集上生成视频流畅且符合指令。
  • 仅需轻量微调,适合快速部署到新任务中。

生成逼真的机器人操作视频是统一具身智能体的感知、规划与行动的重要一步。现有视频扩散模型依赖大量特定领域数据且泛化能力差,而基于语言-图像语料库训练的图像生成模型表现出强组合性,具备合成时间连贯网格图像的能力,暗示其潜在的视频生成能力。我们探索了这类模型在经过轻量级LoRA微调后,能否作为机器人的视觉规划器。提出两阶段框架:(1) 文本条件生成,使用语言指令和初始帧;(2) 轨迹条件生成,使用2D轨迹叠加和相同初始帧。在Jaco Play、Bridge V2和RT1数据集上的实验表明,两种模式均能生成与条件一致的平滑连贯机器人视频。结果表明,预训练图像生成器编码了可迁移的时间先验,可在极少监督下充当类视频机器人规划器。代码已公开。

原文摘要 · Abstract (English)

Generating realistic robotic manipulation videos is an important step toward unifying perception, planning, and action in embodied agents. While existing video diffusion models require large domain-specific datasets and struggle to generalize, recent image generation models trained on language-image corpora exhibit strong compositionality, including the ability to synthesize temporally coherent grid images. This suggests a latent capacity for video-like generation even without explicit temporal modeling. We explore whether such models can serve as visual planners for robots when lightly adapted using LoRA finetuning. We propose a two-part framework that includes: (1) text-conditioned generation, which uses a language instruction and the first frame, and (2) trajectory-conditioned generation, which uses a 2D trajectory overlay and the same initial frame. Experiments on the Jaco Play dataset, Bridge V2, and the RT1 dataset show that both modes produce smooth, coherent robot videos aligned with their respective conditions. Our findings indicate that pretrained image generators encode transferable temporal priors and can function as video-like robotic planners under minimal supervision. Code is released at \href{https://github.com/pangye202264690373/Image-Generation-as-a-Visual-Planner-for-Robotic-Manipulation}{https://github.com/pangye202264690373/Image-Generation-as-a-Visual-Planner-for-Robotic-Manipulation}.

机器人图像生成视觉规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。