arXiv:2503.12042cs.SDcs.CV2025-03CVPR被引 11

提升电影配音质量,让语音情绪与画面同步更精准。

Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing

  • 分两阶段训练:先增强声学建模,再解耦语调与配音风格。
  • 在两个基准上超越现有最佳模型,显著提升情感对齐效果。
  • 适合影视配音、语音合成研究者,尤其关注情绪同步场景。

电影配音需将剧本转化为与视频片段在时间与情感上对齐的语音,同时保留参考音频中的说话人特征。该任务要求模型融合角色表演与复杂语调结构,生成高质量的音视频同步配音。然而,电影配音数据集规模有限且音频背景噪声严重,制约了声学建模性能。为此,本文提出一种声学-语调解耦的两阶段方法,实现高质量配音生成并精确对齐语调。首先,设计语调增强的声学预训练,提升声学建模能力;随后冻结预训练声学系统,构建解耦框架,分别建模语调文本特征与配音风格,同时保持声学质量。此外,引入领域内情感分析模块,缓解不同影片间视觉域偏移的影响,增强情感与语调的一致性。大量实验表明,该方法在两个主要基准上均优于当前最优模型。演示视频见 https://zzdoog.github.io/ProDubber/

原文摘要 · Abstract (English)

Movie dubbing describes the process of transforming a script into speech that aligns temporally and emotionally with a given movie clip while exemplifying the speaker's voice demonstrated in a short reference audio clip. This task demands the model bridge character performances and complicated prosody structures to build a high-quality video-synchronized dubbing track. The limited scale of movie dubbing datasets, along with the background noise inherent in audio data, hinder the acoustic modeling performance of trained models. To address these issues, we propose an acoustic-prosody disentangled two-stage method to achieve high-quality dubbing generation with precise prosody alignment. First, we propose a prosody-enhanced acoustic pre-training to develop robust acoustic modeling capabilities. Then, we freeze the pre-trained acoustic system and design a disentangled framework to model prosodic text features and dubbing style while maintaining acoustic quality. Additionally, we incorporate an in-domain emotion analysis module to reduce the impact of visual domain shifts across different movies, thereby enhancing emotion-prosody alignment. Extensive experiments show that our method performs favorably against the state-of-the-art models on two primary benchmarks. The demos are available at https://zzdoog.github.io/ProDubber/.

语音合成情感对齐电影配音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。