arXiv:2602.00189cs.SDcs.AI2026-02中稿 · Elsevier's \textit…被引 9

用音频生成逼真口型同步的说话头,适配任意人脸。

LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild

  • 基于残差CBAM的U-Net融合音视频信息,提升特征表达能力。
  • 引入语义对齐模块与LPIPS损失,实现精准口型同步和高质量图像生成。
  • 方法通用性强,适用于任意说话人,适合影视合成与虚拟主播场景。

说话头生成领域对音频驱动技术日益关注。核心挑战在于实现唇部动作与音频的高度一致,即口型同步。本文提出一种通用方法 LPIPS-AttnWav2Lip,可根据任意说话人的音频重建其面部图像。采用基于残差CBAM的U-Net架构,更有效地编码与融合音视频模态信息;语义对齐模块扩展生成网络的感受野,高效获取视觉特征的空间与通道信息,并将视觉特征的统计信息与音频潜在向量匹配,实现音频内容信息对视觉信息的调整与注入。为实现精确口型同步并生成高保真图像,方法引入LPIPS损失,模拟人类对图像质量的判断,降低训练过程中的不稳定性。主观与客观评估结果均表明该方法在口型同步精度与视觉质量方面表现优异。代码已开源:https://github.com/FelixChan9527/LPIPS-AttnWav2Lip。

原文摘要 · Abstract (English)

Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchronization. This paper proposes a generic method, LPIPS-AttnWav2Lip, for reconstructing face images of any speaker based on audio. We used the U-Net architecture based on residual CBAM to better encode and fuse audio and visual modal information. Additionally, the semantic alignment module extends the receptive field of the generator network to obtain the spatial and channel information of the visual features efficiently; and match statistical information of visual features with audio latent vector to achieve the adjustment and injection of the audio content information to the visual information. To achieve exact lip synchronization and to generate realistic high-quality images, our approach adopts LPIPS Loss, which simulates human judgment of image quality and reduces instability possibility during the training process. The proposed method achieves outstanding performance in terms of lip synchronization accuracy and visual quality as demonstrated by subjective and objective evaluation results. The code for the paper is available at the following link: https://github.com/FelixChan9527/LPIPS-AttnWav2Lip

说话头生成口型同步音视频对齐U-Net

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。