用生成模型统一解决视频编辑、插入、删除和跟踪任务
Generative Video Propagation
- 通过编码器-生成器框架,从首帧传播修改到整段视频
- 在多个任务上达到领先性能,支持物体形状大幅改变
- 适合需要复杂视频修改的创意与影视制作场景
大规模视频生成模型具有真实模拟自然场景的固有能力。本文证明,通过精心设计的生成式视频传播框架,可统一解决多种视频任务。我们提出的GenProp框架使用选择性内容编码器对原始视频编码,并利用图像到视频生成模型将首帧的修改传播至整段视频。基于实例级视频分割数据集设计数据生成方案,模型通过添加掩码预测解码头并优化区域感知损失,使编码器在保持原始内容的同时,让生成模型有效传播修改区域。该设计实现新可能:在编辑场景中,允许物体形状大幅改变;在插入任务中,插入对象可独立运动;在移除任务中,能有效消除阴影和反光等影响;在跟踪任务中,可同步追踪物体及其关联效应。实验表明,该模型在多种视频任务中表现领先,并提供了深入的框架分析。
原文摘要 · Abstract (English)
Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, GenProp, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object's shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。