用ControlNet让预训练音频模型精准匹配视频,实现音画同步的拟音生成。
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
- 通过频域感知对齐器,将视频时序特征与音频模型特性融合。
- 在基准测试中超越从零训练的基线模型,性能领先。
- 适合想高效部署音画同步生成系统的研究者和开发者。
拟音合成旨在生成在语义和时间上均与视频帧对齐的高质量音频。由于其在创意产业中的广泛应用,该任务日益受到研究界关注。为避免从头训练音频生成模型的复杂性,利用预训练音频生成模型进行视频同步拟音合成成为可行方向。尽管ControlNet可用于添加细粒度控制,但现有方法仅依赖手工可读的时间条件。相比之下,从零训练模型通过预训练视频编码器提取的高维深层特征取得了成功。我们发现基于ControlNet的模型与从零训练模型之间存在性能差距。为此,提出SpecMaskFoley,通过ControlNet引导预训练的SpecMaskGIT模型实现视频同步拟音生成。为释放单个ControlNet分支的潜力,设计频域感知的时间特征对齐器,弥合视频时序特征与预训练模型的时间-频率特性之间的差异,无需以往复杂的条件机制。在通用拟音合成基准上的评估表明,SpecMaskFoley甚至优于强基线的从零训练模型,显著推进了基于ControlNet的拟音合成发展。
原文摘要 · Abstract (English)
Foley synthesis aims to synthesize high-quality audio that is both semantically and temporally aligned with video frames. Given its broad application in creative industries, the task has gained increasing attention in the research community. To avoid the non-trivial task of training audio generative models from scratch, adapting pretrained audio generative models for video-synchronized foley synthesis presents an attractive direction. ControlNet, a method for adding fine-grained controls to pretrained generative models, has been applied to foley synthesis, but its use has been limited to handcrafted human-readable temporal conditions. In contrast, from-scratch models achieved success by leveraging high-dimensional deep features extracted using pretrained video encoders. We have observed a performance gap between ControlNet-based and from-scratch foley models. To narrow this gap, we propose SpecMaskFoley, a method that steers the pretrained SpecMaskGIT model toward video-synchronized foley synthesis via ControlNet. To unlock the potential of a single ControlNet branch, we resolve the discrepancy between the temporal video features and the time-frequency nature of the pretrained SpecMaskGIT via a frequency-aware temporal feature aligner, eliminating the need for complicated conditioning mechanisms widely used in prior arts. Evaluations on a common foley synthesis benchmark demonstrate that SpecMaskFoley could even outperform strong from-scratch baselines, substantially advancing the development of ControlNet-based foley synthesis models. Demo page: https://zzaudio.github.io/SpecMaskFoley_Demo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。