Vera通过分层扩散模型实现视频编辑时内容不变,仅生成改动部分。
Vera: A Layered Diffusion Model for Content-Preserving Video Editing

- 分层生成编辑层与透明度图,分离创意修改与内容保留
- 使用混合注意力机制保证新旧画面融合自然,提升一致性
- 在48.6万帧数据上训练,显著优于现有开源模型的内容保持能力
视频扩散模型在视频生成与编辑中取得显著进展,但内容保持仍是核心挑战:现有方法重绘所有像素,常改变应保持不变的元素(如角色或背景)。我们提出Vera,一种用于内容保持视频编辑的分层扩散框架。Vera不重绘整段视频,而是生成一个编辑层及用于合成的透明度图,从设计上分离创意编辑与内容保留。为促进与源视频的连贯合成,我们将文本到视频DiT扩展为多变换器(MoT)架构,每个层拥有独立的DiT,并通过联合自注意力交互。为支持Vera训练,我们构建了一个高质量分层数据集,包含精确透明度图、多样场景与动态变化及视觉特效。在定量基准与人工偏好评估中,Vera在内容保持方面优于领先开源视频编辑模型,同时在编辑质量上保持竞争力,使用486,000帧分层训练数据。
原文摘要 · Abstract (English)
Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservation remains a core challenge: existing methods regenerate every pixel and often alter elements that should remain unchanged, such as characters or background scenes. We introduce Vera, a layered diffusion framework for content-preserving video editing. Instead of regenerating the entire video, Vera generates an edit layer along with an alpha matte for compositing with the source video, separating creative editing from content preservation by design. To encourage coherent composition with the source video, we extend the text-to-video DiT into a Mixture-of-Transformers (MoT) architecture, with separate DiTs for each layer that interact through joint self-attention. To support the training of Vera, we further construct a high-quality layered dataset with accurate alpha mattes, diverse scenes and dynamics, and visual effects. Across our quantitative benchmark and human preference study, Vera outperforms leading open-source video editing models in content preservation while remaining competitive in edit quality, using 486K frames of layered training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。