arXiv:2501.03059cs.CVcs.AI2025-01CVPR被引 19

用遮罩轨迹显式建模物体运动,让图像转视频更连贯真实。

Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation

  • 引入遮罩运动轨迹作为中间表示,同时捕捉语义与运动信息。
  • 在多物体高动态场景中,视频时序一致性与运动真实性显著提升。
  • 适合需要精准运动控制的视频生成任务,如动画创作与影视预演。

我们研究图像到视频(I2V)生成任务,即根据文本描述将静态图像转换为逼真视频序列。尽管近期方法能生成高质量视觉结果,但在多物体场景下仍难以保证物体运动的准确性和一致性。为此,我们提出一种两阶段组合框架:首先生成显式的中间表示,再基于该表示生成视频。核心创新是引入基于遮罩的运动轨迹作为中间表示,同时编码语义与运动信息,实现对运动与语义的紧凑而丰富的表达。为在第二阶段融入该表示,我们设计了对象级注意力机制:采用空间上每对象的掩码交叉注意力,将物体特定提示注入对应潜在空间区域;并引入掩码时空自注意力,确保每个物体在帧间保持一致。我们在包含多物体和高运动性的挑战性基准上评估,实证表明所提方法在时间连贯性、运动真实性和文本提示忠实度方面达到当前最优。此外,我们提出了新基准 enchmark,涵盖单物体与多物体I2V生成任务,并验证了该方法在此基准上的优势。项目页面见 https://guyyariv.github.io/TTM/。

原文摘要 · Abstract (English)

We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into realistic video sequences based on a textual description. While recent advancements produce photorealistic outputs, they frequently struggle to create videos with accurate and consistent object motion, especially in multi-object scenarios. To address these limitations, we propose a two-stage compositional framework that decomposes I2V generation into: (i) An explicit intermediate representation generation stage, followed by (ii) A video generation stage that is conditioned on this representation. Our key innovation is the introduction of a mask-based motion trajectory as an intermediate representation, that captures both semantic object information and motion, enabling an expressive but compact representation of motion and semantics. To incorporate the learned representation in the second stage, we utilize object-level attention objectives. Specifically, we consider a spatial, per-object, masked-cross attention objective, integrating object-specific prompts into corresponding latent space regions and a masked spatio-temporal self-attention objective, ensuring frame-to-frame consistency for each object. We evaluate our method on challenging benchmarks with multi-object and high-motion scenarios and empirically demonstrate that the proposed method achieves state-of-the-art results in temporal coherence, motion realism, and text-prompt faithfulness. Additionally, we introduce \benchmark, a new challenging benchmark for single-object and multi-object I2V generation, and demonstrate our method's superiority on this benchmark. Project page is available at https://guyyariv.github.io/TTM/.

图像转视频运动建模视频生成注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。