通过分层合成提升视频生成的控制力与真实感
Layer-Aware Video Composition via Split-then-Merge
- 将视频拆分为动态前景与背景,自动生成组合样本
- 在多个评估中超越现有最佳方法,提升真实感与一致性
- 适合需要精细控制视频内容生成的研究者与开发者
我们提出 Split-then-Merge(StM)框架,旨在提升生成式视频组合的控制能力并缓解数据稀缺问题。与依赖标注数据或人工规则的传统方法不同,StM将大规模未标注视频拆分为动态前景与背景层,并自动生成组合样本,以学习动态主体与多样化场景间的复杂交互关系。该过程使模型掌握真实视频生成所需的复合动态机制。StM引入一种新型变换感知训练流程,结合多层融合与增强技术,实现功能感知的组合;同时采用身份保持损失,在融合过程中确保前景内容的完整性。实验表明,StM在定量指标和人类及视觉语言模型(VLM)的定性评估中均优于当前最佳方法。
原文摘要 · Abstract (English)
We present Split-then-Merge (StM), a novel framework designed to enhance control in generative video composition and address its data scarcity problem. Unlike conventional methods relying on annotated datasets or handcrafted rules, StM splits a large corpus of unlabeled videos into dynamic foreground and background layers, then self-composes them to learn how dynamic subjects interact with diverse scenes. This process enables the model to learn the complex compositional dynamics required for realistic video generation. StM introduces a novel transformation-aware training pipeline that utilizes a multi-layer fusion and augmentation to achieve affordance-aware composition, alongside an identity-preservation loss that maintains foreground fidelity during blending. Experiments show StM outperforms SoTA methods in both quantitative benchmarks and in humans/VLM-based qualitative evaluations. More details are available at our project page: https://split-then-merge.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。