arXiv:2503.04067cs.CV2025-03中稿 · ICMR 2025被引 1

从频域出发,实现高保真实时人脸说话视频生成

FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis

  • 在频域建模视觉与音频特征,提升合成自然度
  • 支持高清实时生成,唇动同步精度超越现有方法
  • 可灵活切换单张图或视频配音,适用场景广

在音频驱动的人脸说话视频生成中,保持高保真唇语同步仍具挑战。现有方法多基于像素域,虽有高质结果但计算成本高;部分方法资源消耗低但唇动错位。我们发现合成视频与真实视频在频域存在显著差异,而当前研究未考虑此方面。为此,提出频率调制的高保真实时人脸说话视频合成框架FREAK,从频域视角建模,增强合成质量。FREAK引入两个新型频域模块:视觉编码频调制器(VEFM)在频域耦合多尺度视觉特征,更好保留视觉频域信息,缩小合成帧与真实帧的频谱差距;音频视觉频调制器(AVFM)帮助模型学习频域中的说话模式,提升音画同步。同时在像素域与频域联合优化。此外,支持单次输入与视频配音无缝切换,兼具灵活性。实验表明,FREAK可实现实时高清生成,细节丰富、唇动精准,优于现有最优方法。

原文摘要 · Abstract (English)

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computational costs. Some approaches perform well on specific individuals with low resources, yet still exhibit mismatched lip movements. The aforementioned methods are modeled in the pixel domain. We observed that there are noticeable discrepancies in the frequency domain between the synthesized talking videos and natural videos. Currently, no research on talking portrait synthesis has considered this aspect. To address this, we propose a FREquency-modulated, high-fidelity, and real-time Audio-driven talKing portrait synthesis framework, named FREAK, which models talking portraits from the frequency domain perspective, enhancing the fidelity and naturalness of the synthesized portraits. FREAK introduces two novel frequency-based modules: 1) the Visual Encoding Frequency Modulator (VEFM) to couple multi-scale visual features in the frequency domain, better preserving visual frequency information and reducing the gap in the frequency spectrum between synthesized and natural frames. and 2) the Audio Visual Frequency Modulator (AVFM) to help the model learn the talking pattern in the frequency domain and improve audio-visual synchronization. Additionally, we optimize the model in both pixel domain and frequency domain jointly. Furthermore, FREAK supports seamless switching between one-shot and video dubbing settings, offering enhanced flexibility. Due to its superior performance, it can simultaneously support high-resolution video results and real-time inference. Extensive experiments demonstrate that our method synthesizes high-fidelity talking portraits with detailed facial textures and precise lip synchronization in real-time, outperforming state-of-the-art methods.

人脸说话频域建模实时生成音画同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。