用目标图像引导视频生成,让机器人提前想象操作结果。
Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion
- 以目标图像为约束,生成符合物理规律的视觉规划轨迹。
- 在物体操作和图像编辑任务中,目标对齐率提升23%,空间一致性更强。
- 适合需要精准动作规划的机器人系统,尤其擅长复杂场景下的任务预演。
具身视觉规划旨在通过想象场景向期望目标演化的过程来指导操作。视频扩散模型凭借其图像到视频的生成能力,为这种视觉想象提供了良好基础。然而,现有方法多为前向预测,仅基于初始观测生成轨迹,缺乏显式的目标建模,常导致空间漂移和目标偏离。为此,我们提出Envision,一种基于扩散模型的具身视觉规划框架。该框架分两阶段运行:首先,目标意象模型识别任务相关区域,通过场景与指令间的区域感知交叉注意力,合成一个连贯的目标图像;其次,基于首尾帧条件的视频扩散模型(FL2V),环境-目标视频模型在初始观测与目标图像间插值,生成平滑且物理合理的视频轨迹。在物体操作与图像编辑基准测试中,Envision在目标对齐、空间一致性和物体保留方面均优于基线。生成的视觉计划可直接用于下游机器人规划与控制,为具身智能体提供可靠指导。
原文摘要 · Abstract (English)
Embodied visual planning aims to enable manipulation tasks by imagining how a scene evolves toward a desired goal and using the imagined trajectories to guide actions. Video diffusion models, through their image-to-video generation capability, provide a promising foundation for such visual imagination. However, existing approaches are largely forward predictive, generating trajectories conditioned on the initial observation without explicit goal modeling, thus often leading to spatial drift and goal misalignment. To address these challenges, we propose Envision, a diffusion-based framework that performs visual planning for embodied agents. By explicitly constraining the generation with a goal image, our method enforces physical plausibility and goal consistency throughout the generated trajectory. Specifically, Envision operates in two stages. First, a Goal Imagery Model identifies task-relevant regions, performs region-aware cross attention between the scene and the instruction, and synthesizes a coherent goal image that captures the desired outcome. Then, an Env-Goal Video Model, built upon a first-and-last-frame-conditioned video diffusion model (FL2V), interpolates between the initial observation and the goal image, producing smooth and physically plausible video trajectories that connect the start and goal states. Experiments on object manipulation and image editing benchmarks demonstrate that Envision achieves superior goal alignment, spatial consistency, and object preservation compared to baselines. The resulting visual plans can directly support downstream robotic planning and control, providing reliable guidance for embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。