arXiv:2511.17839cs.CV2025-11中稿 · WACV 2026

统一图像与视频生成,让模型更懂动作意图与时间演化。

Show Me: Unifying Instructional Image and Video Generation with Diffusion Models

  • 通过激活扩散模型的时空组件,统一处理图像与视频生成任务。
  • 在多个基准上超越专家模型,图像编辑更真实,视频预测更连贯。
  • 适合构建交互式世界模拟器,尤其关注动作目标与时间演变。

在特定上下文中生成视觉化操作指令,对开发交互式世界模拟器至关重要。以往工作分别采用文本引导图像编辑或视频预测,但二者常被孤立处理。这暴露了根本问题:图像编辑忽略动作的时间演进,而视频预测常忽视预期结果。为此,我们提出ShowMe,一个统一框架,通过选择性激活视频扩散模型的空间与时间组件,同时实现两项任务。此外,引入结构与运动一致性奖励,提升结构保真度与时间连贯性。该统一带来双重优势:视频预训练获得的空间知识增强了非刚性图像编辑的上下文一致性和真实性;指令引导的编辑阶段则赋予模型更强的目标导向推理能力,用于视频预测。在多个基准上的实验表明,我们的方法在视觉指令图像与视频生成上均优于现有专家模型,凸显视频扩散模型作为统一的动作-物体状态转换器的潜力。

原文摘要 · Abstract (English)

Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image manipulation or video prediction, these tasks are typically treated in isolation. This separation reveals a fundamental issue: image manipulation methods overlook how actions unfold over time, while video prediction models often ignore the intended outcomes. To this end, we propose ShowMe, a unified framework that enables both tasks by selectively activating the spatial and temporal components of video diffusion models. In addition, we introduce structure and motion consistency rewards to improve structural fidelity and temporal coherence. Notably, this unification brings dual benefits: the spatial knowledge gained through video pretraining enhances contextual consistency and realism in non-rigid image edits, while the instruction-guided manipulation stage equips the model with stronger goal-oriented reasoning for video prediction. Experiments on diverse benchmarks demonstrate that our method outperforms expert models in both instructional image and video generation, highlighting the strength of video diffusion models as a unified action-object state transformer.

视频生成扩散模型指令生成统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。