用双编码器融合音频内容与动态,提升说话头视频的口型同步和画质。
Dual Audio-Centric Modality Coupling for Talking Head Generation
- 双编码器分别捕捉音频语义与动态同步信息
- 在多个数据集上口型同步准确率优于现有方法
- 适合虚拟主播、数字人等需要高保真口型同步的应用
基于音频驱动的说话头视频生成是计算机视觉与图形学中的关键挑战,广泛应用于虚拟形象和数字媒体。传统方法难以捕捉音频与面部动作间的复杂互动,导致口型不同步和画面质量不佳。本文提出一种基于NeRF的新框架——双音频中心模态耦合(DAMC),有效融合音频输入的内容与动态特征。通过双编码器结构,内容感知编码器提取语义内容,动态同步编码器确保视觉精准对齐。二者通过交叉同步融合模块(CSFM)融合,增强表征能力与口型同步效果。大量实验表明,该方法在口型同步精度与图像质量等关键指标上超越现有最优方法,且在各类音频输入(包括文本转语音系统生成的合成语音)下均表现稳健。结果为高质量音频驱动说话头生成提供了可行方案,并展示了可扩展性。
原文摘要 · Abstract (English)
The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the complex interaction between audio and facial dynamics, leading to lip synchronization and visual quality issues. In this paper, we propose a novel NeRF-based framework, Dual Audio-Centric Modality Coupling (DAMC), which effectively integrates content and dynamic features from audio inputs. By leveraging a dual encoder structure, DAMC captures semantic content through the Content-Aware Encoder and ensures precise visual synchronization through the Dynamic-Sync Encoder. These features are fused using a Cross-Synchronized Fusion Module (CSFM), enhancing content representation and lip synchronization. Extensive experiments show that our method outperforms existing state-of-the-art approaches in key metrics such as lip synchronization accuracy and image quality, demonstrating robust generalization across various audio inputs, including synthetic speech from text-to-speech (TTS) systems. Our results provide a promising solution for high-quality, audio-driven talking head generation and present a scalable approach for creating realistic talking heads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。