通过区域约束提升指令视频编辑的精准度与一致性。
Region-Constraint In-Context Generation for Instructional Video Editing
- 联合编码源目标视频,用潜空间和注意力正则化控制编辑区域
- 在4个任务上实现显著优于基线的编辑精度与视觉质量
- 适合需要高精度区域控制的视频编辑研究者
近期的上下文生成范式在指令图像编辑中展现出优异的数据效率与合成质量。然而,将其应用于基于指令的视频编辑仍具挑战:未明确指定编辑区域会导致区域不准及去噪过程中编辑区与非编辑区的令牌干扰。为此,我们提出ReCo,一种新型指令视频编辑范式,首次在上下文生成中建模编辑区与非编辑区间的约束关系。技术上,ReCo横向拼接源视频与目标视频进行联合去噪,并引入潜空间与注意力正则化项,分别作用于单步反向去噪的潜在表示与注意力图。前者增强编辑区在源/目标间的潜在差异,降低非编辑区差异,强化编辑区修改并减少非编辑区意外内容生成;后者抑制编辑区对源视频对应区域的注意力,减轻新物体生成时的干扰。此外,我们构建了大规模高质量视频编辑数据集ReCo-Data,包含50万条指令-视频配对,用于模型训练。在四个主流指令视频编辑任务上的实验表明,所提方法具有显著优势。
原文摘要 · Abstract (English)
The In-context generation paradigm recently has demonstrated strong power in instructional image editing with both data efficiency and synthesis quality. Nevertheless, shaping such in-context learning for instruction-based video editing is not trivial. Without specifying editing regions, the results can suffer from the problem of inaccurate editing regions and the token interference between editing and non-editing areas during denoising. To address these, we present ReCo, a new instructional video editing paradigm that novelly delves into constraint modeling between editing and non-editing regions during in-context generation. Technically, ReCo width-wise concatenates source and target video for joint denoising. To calibrate video diffusion learning, ReCo capitalizes on two regularization terms, i.e., latent and attention regularization, conducting on one-step backward denoised latents and attention maps, respectively. The former increases the latent discrepancy of the editing region between source and target videos while reducing that of non-editing areas, emphasizing the modification on editing area and alleviating outside unexpected content generation. The latter suppresses the attention of tokens in the editing region to the tokens in counterpart of the source video, thereby mitigating their interference during novel object generation in target video. Furthermore, we propose a large-scale, high-quality video editing dataset, i.e., ReCo-Data, comprising 500K instruction-video pairs to benefit model training. Extensive experiments conducted on four major instruction-based video editing tasks demonstrate the superiority of our proposal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。