提出显式分层视频建模方法,提升物体插入与图层分解效果
Explicit Layer Modeling for Video Object Insertion and Layer Decomposition

- 构建三元组数据集TriLayer,提供前景与背景对齐标注
- 设计双分支扩散模型DBL-Diffusion,联合建模彩色合成与透明图层
- 支持真实合成与灵活后期编辑,适合视频编辑与生成任务
现有视频编辑系统普遍缺乏显式的分层视频表示,限制了真实合成、物体复用和一致操作的能力。这一问题在视频物体插入与视频图层分解中尤为突出,因缺乏显式前景层监督,现有方法依赖隐式推理或逐场景优化。本文提出TriLayer,一个大规模三元组视频数据集,包含对齐的合成视频、背景视频和前景视频,其中前景层包含物体外观及关联视觉特效。该显式监督使模型可直接学习分层视频表示,而非隐式推断。基于此数据集,我们提出DBL-Diffusion,一种双分支扩散框架,通过共享去噪过程与跨分支交互,联合建模RGB合成视频与RGBA前景层。在两个任务中实现:DBL-Insert用于分层物体插入,生成显式RGBA层以支持真实合成与灵活后处理;DBL-Decompose用于视频图层分解,利用三元组监督恢复前景与背景层。实验表明,显式分层建模显著提升插入保真度与分解质量。
原文摘要 · Abstract (English)
Most video editing systems still lack explicit layered video representations, limiting their ability to perform realistic compositing, object reuse, and consistent manipulation. This limitation is especially pronounced in video object insertion and video layer decomposition, where existing methods rely on implicit inference or per-scene optimization due to the absence of explicit foreground-layer supervision. We introduce TriLayer, a large-scale triplet video dataset containing aligned composite, background, and foreground videos, where the foreground layers include both object appearance and associated visual effects. This explicit supervision enables models to learn layered video representations directly rather than inferring them implicitly. Building on this dataset, we propose DBL-Diffusion, a dual-branch diffusion framework that jointly models RGB composites and RGBA foreground layers through shared denoising and cross-branch interaction. We instantiate the framework in two tasks: DBL-Insert for layered object insertion, which generates explicit RGBA layers for realistic compositing and flexible post-editing, and DBL-Decompose for video layer decomposition, which recovers foreground and background layers using triplet supervision. Experiments demonstrate that explicit layer modeling substantially improves both insertion fidelity and decomposition quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。