用视觉增强的TTS模型实现跨语言视频配音,同步更准、语音更自然。
SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
- 基于预训练TTS模型,融合视觉信息微调提升音画一致
- 跨语言合成中抑制语种干扰,实现高质量多语言配音
- 适合影视翻译、多语种内容生成场景
视频配音旨在生成与视觉内容精确时序对齐的高保真语音。现有方法在语音自然度和音画同步方面仍有不足,且仅限于单语场景。为此,我们提出SyncVoice,一个基于预训练文本转语音(TTS)模型的视觉增强视频配音框架。通过在音视频数据上微调TTS模型,实现强音视频一致性。设计双说话人编码器以有效缓解跨语言语音合成中的语种干扰,并探索其在视频翻译场景中的应用。实验结果表明,SyncVoice在生成高保真语音的同时具备优异的同步性能,展现出在视频配音任务中的潜力。
原文摘要 · Abstract (English)
Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to monolingual settings. To address these challenges, we propose SyncVoice, a vision-augmented video dubbing framework built upon a pretrained text-to-speech (TTS) model. By fine-tuning the TTS model on audio-visual data, we achieve strong audiovisual consistency. We propose a Dual Speaker Encoder to effectively mitigate inter-language interference in cross-lingual speech synthesis and explore the application of video dubbing in video translation scenarios. Experimental results show that SyncVoice achieves high-fidelity speech generation with strong synchronization performance, demonstrating its potential in video dubbing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。