arXiv:2505.01263cs.MMcs.CV2025-05被引 9

用大模型提升配音同步与音质,让字幕说话更自然

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing

  • 用大语言模型学习剧本与参考音频的语义,实现语音对齐
  • 双对比对齐减少发音混淆,提升唇动同步精度
  • 基于流匹配的声学增强,保留人声特征并提升清晰度

电影配音旨在将剧本转化为与视频片段在时间与情感上一致、同时保持参考音频声纹特征的语音。现有方法多关注降低词错误率,忽视唇形同步与声学质量。为此,我们提出基于大语言模型(LLM)与流匹配的配音框架 FlowDubber。首先采用 Qwen2.5 作为骨干模型,从剧本与参考音频中学习上下文序列;其次引入语义感知学习,在音素层面捕捉语言模型的语义知识;再通过双对比对齐(DCA)增强语音与唇动的互对齐,减少相似音素混淆;最后提出基于流的声学增强(FVE),利用大模型引导声学流匹配以提升语音清晰度,并结合仿射风格先验,在从噪声重建梅尔谱图时增强身份一致性。大量实验表明,该方法在两个主流基准上均优于多个前沿方法。

原文摘要 · Abstract (English)

Movie Dubbing aims to convert scripts into speeches that align with the given movie clip in both temporal and emotional aspects while preserving the vocal timbre of a given brief reference audio. Existing methods focus primarily on reducing the word error rate while ignoring the importance of lip-sync and acoustic quality. To address these issues, we propose a large language model (LLM) based flow matching architecture for dubbing, named FlowDubber, which achieves high-quality audio-visual sync and pronunciation by incorporating a large speech language model and dual contrastive aligning while achieving better acoustic quality via the proposed voice-enhanced flow matching than previous works. First, we introduce Qwen2.5 as the backbone of LLM to learn the in-context sequence from movie scripts and reference audio. Then, the proposed semantic-aware learning focuses on capturing LLM semantic knowledge at the phoneme level. Next, dual contrastive aligning (DCA) boosts mutual alignment with lip movement, reducing ambiguities where similar phonemes might be confused. Finally, the proposed Flow-based Voice Enhancing (FVE) improves acoustic quality in two aspects, which introduces an LLM-based acoustics flow matching guidance to strengthen clarity and uses affine style prior to enhance identity when recovering noise into mel-spectrograms via gradient vector field prediction. Extensive experiments demonstrate that our method outperforms several state-of-the-art methods on two primary benchmarks.

语音生成大模型配音流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。