arXiv:2410.00990cs.CV2024-10被引 1

用数学理论提升语音驱动人脸生成的纹理清晰度与抗噪能力

Lipschitz-Driven Noise Robustness in VQ-AE for High-Frequency Texture Repair in ID-Specific Talking Heads

  • 基于利普希茨连续性理论,发现VQ-AE有抗噪上限
  • 仅需训练身份专属离散表示,就可实现高效去噪
  • 插件式设计,适合工业级实时语音驱动人脸生成

语音驱动的身份特定说话头生成在影视和虚拟现实应用中前景广阔。现有方法多为端到端架构,虽取得进展,但受限于模型容量,难以捕捉高频纹理。为此,我们提出一种新颖的后处理框架,基于神经网络的利普希茨连续性理论,证明了向量量化自编码器(VQ-AE)具备关键的噪声容忍特性,并建立了噪声鲁棒性上界(NRoUB)。该理论表明,仅通过训练身份专属的神经离散表示即可获得身份特定的去噪器,无需额外网络。基于此,我们提出可插拔的时空优化型VQ-AE(SOVQAE),增强其NRoUB,实现时序一致的去噪效果。为便于部署,进一步构建融合预训练Wav2Lip与SOVQAE的级联流水线,完成身份特定说话头生成。实验表明,该流水线在视频质量与分布外唇音同步鲁棒性方面达到当前最优表现,仅需数小时消费级GPU训练时间,且支持实时运行,兼具高效性与实用性。

原文摘要 · Abstract (English)

Audio-driven IDentity-specific Talking Head Generation (ID-specific THG) has shown increasing promise for applications in filmmaking and virtual reality. Existing approaches are generally constructed as end-to-end paradigms, and have achieved significant progress. However, they often struggle to capture high-frequency textures due to limited model capacity. To address these limitations, we adopt a simple yet efficient post-processing framework -- unlike previous studies that focus solely on end-to-end training -- guided by our theoretical insights. Specifically, leveraging the \textit{Lipschitz Continuity Theory} of neural networks, we prove a crucial noise tolerance property for the Vector Quantized AutoEncoder (VQ-AE), and establish the existence of a Noise Robustness Upper Bound (NRoUB). This insight reveals that we can efficiently obtain an identity-specific denoiser by training an identity-specific neural discrete representation, without requiring an extra network. Based on this theoretical foundation, we propose a plug-and-play Space-Optimized VQ-AE (SOVQAE) with enhanced NRoUB to achieve temporally-consistent denoising. For practical deployment, we further introduce a cascade pipeline combining a pretrained Wav2Lip model with SOVQAE to perform ID-specific THG. Our experiments demonstrate that this pipeline achieves \textit{state-of-the-art} performance in video quality and robustness for out-of-distribution lip synchronization, surpassing existing identity-specific THG methods. In addition, the pipeline requires only a couple of consumer GPU hours and runs in real time, which is both efficient and practical for industry applications.

说话头生成VQ-AE抗噪能力实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。