arXiv:2412.03430cs.CVcs.LG2024-12被引 2

用多尺度频谱扩散模型生成逼真歌唱视频,效果超越现有方法。

SINGER: Vivid Audio-driven Singing Video Generation with Multi-scale Spectral Diffusion Model

  • 设计多尺度频谱模块,学习歌唱的频域特征
  • 引入频谱过滤模块,捕捉歌唱时的人体动作规律
  • 构建真实世界歌唱数据集,推动领域发展

近年来生成模型在说话人脸视频生成方面取得显著进展,但歌唱视频生成仍鲜有研究。由于说话与歌唱在音频特征和行为表现上的根本差异,现有说话人脸生成模型在歌唱任务上表现不佳。我们发现歌唱与说话音频在频率和振幅上存在显著差异。为此,我们设计了多尺度频谱模块,帮助模型在频域学习歌唱模式;同时开发频谱过滤模块,使模型能捕捉与歌唱音频相关的人体行为。这两个模块被集成到扩散模型中,形成SINGER模型。此外,高质量真实场景歌唱视频数据的缺失阻碍了该领域发展。为此,我们收集了一个真实环境下的音视频歌唱数据集,以支持后续研究。实验表明,SINGER能够生成生动逼真的歌唱视频,在客观和主观评估中均优于现有最先进方法。

原文摘要 · Abstract (English)

Recent advancements in generative models have significantly enhanced talking face video generation, yet singing video generation remains underexplored. The differences between human talking and singing limit the performance of existing talking face video generation models when applied to singing. The fundamental differences between talking and singing-specifically in audio characteristics and behavioral expressions-limit the effectiveness of existing models. We observe that the differences between singing and talking audios manifest in terms of frequency and amplitude. To address this, we have designed a multi-scale spectral module to help the model learn singing patterns in the spectral domain. Additionally, we develop a spectral-filtering module that aids the model in learning the human behaviors associated with singing audio. These two modules are integrated into the diffusion model to enhance singing video generation performance, resulting in our proposed model, SINGER. Furthermore, the lack of high-quality real-world singing face videos has hindered the development of the singing video generation community. To address this gap, we have collected an in-the-wild audio-visual singing dataset to facilitate research in this area. Our experiments demonstrate that SINGER is capable of generating vivid singing videos and outperforms state-of-the-art methods in both objective and subjective evaluations.

歌唱生成扩散模型音视频同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。