通过音素与光流一致性提升语音驱动唇动的自然流畅度
FluentLip: A Phonemes-Based Two-stage Approach for Audio-Driven Lip Synthesis with Optical Flow Consistency
- 分两阶段建模,融合音素信息增强同步性
- 光流一致性损失使帧间过渡更自然,PER降低35.2%
- 结合扩散链训练GAN,提升稳定性和生成效率
语音驱动唇动生成连续唇部动作图像是一项挑战性任务。尽管已有研究在同步性和视觉质量上取得进展,但唇部可理解性与视频流畅性仍是难题。本文提出FluentLip,一种基于音素的两阶段音频驱动唇动合成方法,包含三项核心策略:引入音素提取器与编码器,融合音频与音素信息以促进多模态学习;采用光流一致性损失,确保帧间过渡自然;在生成对抗网络(GAN)训练中引入扩散链,提升训练稳定性与效率。通过五项指标(含新提出的音素错误率PER)对比五种前沿方法,实验表明,FluentLip在平滑性与自然度上表现优异,相比SOTA方法,FID提升16.3%,PER降低35.2%。
原文摘要 · Abstract (English)
Generating consecutive images of lip movements that align with a given speech in audio-driven lip synthesis is a challenging task. While previous studies have made strides in synchronization and visual quality, lip intelligibility and video fluency remain persistent challenges. This work proposes FluentLip, a two-stage approach for audio-driven lip synthesis, incorporating three featured strategies. To improve lip synchronization and intelligibility, we integrate a phoneme extractor and encoder to generate a fusion of audio and phoneme information for multimodal learning. Additionally, we employ optical flow consistency loss to ensure natural transitions between image frames. Furthermore, we incorporate a diffusion chain during the training of Generative Adversarial Networks (GANs) to improve both stability and efficiency. We evaluate our proposed FluentLip through extensive experiments, comparing it with five state-of-the-art (SOTA) approaches across five metrics, including a proposed metric called Phoneme Error Rate (PER) that evaluates lip pose intelligibility and video fluency. The experimental results demonstrate that our FluentLip approach is highly competitive, achieving significant improvements in smoothness and naturalness. In particular, it outperforms these SOTA approaches by approximately $\textbf{16.3%}$ in Fréchet Inception Distance (FID) and $\textbf{35.2%}$ in PER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。