用人脸动作生成同步配音,语音自然流畅。
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
- 融合视频特征的神经编解码语言模型,实现音画同步。
- 在真实场景数据集上,口型对齐准确率超基线方法。
- 适合影视制作与无障碍语音辅助应用。
我们提出 VoiceCraft-Dub,一种自动化视频配音的新方法,可从文本和面部动作生成高质量语音。该任务在影视制作、多媒体创作及辅助失语者方面有广泛应用。基于神经编解码语言模型(NCLM)在语音合成中的成功,本方法通过引入视频特征,使合成语音在时间上同步且情感表达与面部动作一致,同时保持自然语调。为融合视觉信息,我们设计适配器将面部特征映射至 NCLM 的词元空间,并引入音视频融合层,在 NCLM 框架内整合多模态信息。此外,我们构建了 CelebV-Dub,一个专为自动化视频配音设计的富有表现力的真实世界视频数据集。大量实验表明,该模型在人耳感知评估中表现优异,客观评价也优于现有方法,实现了高保真、可懂且自然的语音合成,口型同步精准。我们还将其拓展至视频到语音任务,验证了其多功能性。
原文摘要 · Abstract (English)
We present VoiceCraft-Dub, a novel approach for automated video dubbing that synthesizes high-quality speech from text and facial cues. This task has broad applications in filmmaking, multimedia creation, and assisting voice-impaired individuals. Building on the success of Neural Codec Language Models (NCLMs) for speech synthesis, our method extends their capabilities by incorporating video features, ensuring that synthesized speech is time-synchronized and expressively aligned with facial movements while preserving natural prosody. To inject visual cues, we design adapters to align facial features with the NCLM token space and introduce audio-visual fusion layers to merge audio-visual information within the NCLM framework. Additionally, we curate CelebV-Dub, a new dataset of expressive, real-world videos specifically designed for automated video dubbing. Extensive experiments show that our model achieves high-quality, intelligible, and natural speech synthesis with accurate lip synchronization, outperforming existing methods in human perception and performing favorably in objective evaluations. We also adapt VoiceCraft-Dub for the video-to-speech task, demonstrating its versatility for various applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。