arXiv:2508.04228cs.CVcs.AI2025-08被引 2

一次性生成可编辑的分层视频,提升专业视频制作效率。

LayerT2V: A Unified Multi-Layer Video Generation Framework

  • 通过时序压缩共享生成轨迹,统一建模多层视频输出。
  • 在多个数据集上实现更高画质、更稳时序与更强层间一致性。
  • 适合需要分层编辑的影视、广告等专业视频创作场景。

文本到视频生成虽进展迅速,但现有方法通常仅输出最终合成视频,缺乏可编辑的分层表示,限制了其在专业流程中的应用。我们提出 LayerT2V,一个统一的多层视频生成框架,可在单次推理中生成完整视频、独立背景层及多个带对应透明度图的前景RGB层。关键洞察是,近期视频生成主干在时空上均采用高密度压缩,使我们能将多层表示沿时间维度序列化,并在共享生成轨迹上联合建模,将层间一致性自然转化为内在目标,从而提升语义对齐与时间连贯性。为缓解层间歧义与条件泄露,我们在共享DiT主干上引入LayerAdaLN和层感知交叉注意力调制。LayerT2V经三阶段训练:α掩码VAE适配、联合多层学习、多前景扩展。我们还提出了首个大规模多层视频生成数据集VidLayer。大量实验表明,LayerT2V在视觉保真度、时间一致性和跨层一致性上显著优于先前方法。

原文摘要 · Abstract (English)

Text-to-video generation has advanced rapidly, but existing methods typically output only the final composited video and lack editable layered representations, limiting their use in professional workflows. We propose \textbf{LayerT2V}, a unified multi-layer video generation framework that produces multiple semantically consistent outputs in a single inference pass: the full video, an independent background layer, and multiple foreground RGB layers with corresponding alpha mattes. Our key insight is that recent video generation backbones use high compression in both time and space, enabling us to serialize multiple layer representations along the temporal dimension and jointly model them on a shared generation trajectory. This turns cross-layer consistency into an intrinsic objective, improving semantic alignment and temporal coherence. To mitigate layer ambiguity and conditional leakage, we augment a shared DiT backbone with LayerAdaLN and layer-aware cross-attention modulation. LayerT2V is trained in three stages: alpha mask VAE adaptation, joint multi-layer learning, and multi-foreground extension. We also introduce \textbf{VidLayer}, the first large-scale dataset for multi-layer video generation. Extensive experiments demonstrate that LayerT2V substantially outperforms prior methods in visual fidelity, temporal consistency, and cross-layer coherence.

视频生成分层编辑扩散模型多层建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。