arXiv:2409.02634cs.CV2024-09ICLR被引 113

用纯音频生成自然口型动作,无需额外运动模板。

Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

  • 设计时空模块,捕捉长期动作依赖关系。
  • 在多个数据集上超越现有方法,动作更自然流畅。
  • 适合做语音驱动虚拟人、数字人动画的开发者。

随着基于扩散模型的视频生成技术发展,语音驱动的人像视频生成在动作自然性和面部细节还原方面取得显著进展。然而,现有方法受限于语音信号对动作的控制能力,常需引入辅助空间信号以稳定动作,这可能损害动作的自然性与自由度。本文提出一种端到端的纯音频条件视频扩散模型 Loopy,设计了跨片段与片内时间模块,以及音频到潜在表示模块,使模型能从数据中学习长期动作模式,增强语音与人脸动作的相关性。该方法无需人工设定空间运动模板即可在推理时稳定生成。大量实验表明,Loopy 在多种场景下均优于当前主流语音驱动人脸扩散模型,生成结果更逼真、质量更高。

原文摘要 · Abstract (English)

With the introduction of diffusion-based video generation techniques, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to the limited control of audio signals in driving human motion, existing methods often add auxiliary spatial signals to stabilize movements, which may compromise the naturalness and freedom of motion. In this paper, we propose an end-to-end audio-only conditioned video diffusion model named Loopy. Specifically, we designed an inter- and intra-clip temporal module and an audio-to-latents module, enabling the model to leverage long-term motion information from the data to learn natural motion patterns and improving audio-portrait movement correlation. This method removes the need for manually specified spatial motion templates used in existing methods to constrain motion during inference. Extensive experiments show that Loopy outperforms recent audio-driven portrait diffusion models, delivering more lifelike and high-quality results across various scenarios.

语音生成扩散模型人脸动画

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。