arXiv:2505.16279cs.MMcs.CV2025-05中稿 · Interspeech 2025

用多模态模型提升电影配音的风格适配与细节表现

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

  • 结合视觉语言模型识别配音类型与说话人特征
  • 在多个数据集上实现最高19.08%的情感相似度提升
  • 适合需要高质量跨风格配音的应用场景

当前电影配音技术能通过参考音色和输入视频生成同步且情感准确的语音,但对配音风格适应、对话/独白处理及说话人年龄性别等细微特征的考虑仍不足。为此,我们提出一种多模态生成框架:首先利用多模态大视觉语言模型(VLM)分析视觉输入,识别配音类型与细粒度属性;其次基于多模态输入,使用大语音生成模型生成高质量配音。此外,构建了一个带配音类型与细微特征标注的电影配音数据集,以增强对电影的理解并提升配音质量。在多个基准数据集上的实验表明,该框架优于现有最先进方法,具体在LSE-D、SPK-SIM、EMO-SIM和MCD指标上分别提升1.09%、8.80%、19.08%和18.74%。

原文摘要 · Abstract (English)

Current movie dubbing technology can produce the desired speech using a reference voice and input video, maintaining perfect synchronization with the visuals while effectively conveying the intended emotions. However, crucial aspects of movie dubbing, including adaptation to various dubbing styles, effective handling of dialogue, narration, and monologues, as well as consideration of subtle details such as speaker age and gender, remain insufficiently explored. To tackle these challenges, we introduce a multi-modal generative framework. First, it utilizes a multi-modal large vision-language model (VLM) to analyze visual inputs, enabling the recognition of dubbing types and fine-grained attributes. Second, it produces high-quality dubbing using large speech generation models, guided by multi-modal inputs. Additionally, a movie dubbing dataset with annotations for dubbing types and subtle details is constructed to enhance movie understanding and improve dubbing quality for the proposed multi-modal framework. Experimental results across multiple benchmark datasets show superior performance compared to state-of-the-art (SOTA) methods. In details, the LSE-D, SPK-SIM, EMO-SIM, and MCD exhibit improvements of up to 1.09%, 8.80%, 19.08%, and 18.74%, respectively.

多模态语音生成电影配音视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。