arXiv:2503.23660cs.CV2025-03被引 7

用多模态推理让配音更贴合角色和场景,提升质量与风格适配性。

DeepDubber-V1: Towards High Quality and Dialogue, Narration, Monologue Adaptive Movie Dubbing Via Multi-Modal Chain-of-Thoughts Reasoning Guidance

  • 通过视觉信息进行多模态思维链推理,理解角色年龄、性别与配音风格。
  • 在多个数据集上,语音相似度提升至89.74%,错误率降低至23.20%。
  • 适合需要高质量、风格化电影配音的创作者与影视工业化生产者。

当前电影配音技术可从语音提示生成匹配画面的语音,实现音画同步与情绪传达。然而,在配音风格适配、对话/旁白/独白处理以及说话人年龄、性别等细微特征理解方面仍缺乏深入研究。为此,我们提出一种多模态大语言模型框架:首先利用多模态思维链(CoT)推理方法分析视觉输入,理解配音风格与细粒度属性;其次通过大语音生成模型,在多模态条件引导下生成高质量配音。我们还构建了一个带有思维链标注的电影配音数据集。评估结果表明,相比现有方法,该框架在多个数据集上性能显著提升。具体而言,在V2C Animation数据集上,双语语音相似度(SPK-SIM)从82.48%升至89.74%,情感相似度(EMO-SIM)从66.24%增至78.88%;在Grid数据集上,语音失真率(LSE-D)从14.79降至14.63,音乐-语音一致性(MCD-SL)从5.24降至4.74;在自建的CoT-Movie-Dubbing数据集初始推理设置中,SPK-SIM从64.03%升至83.42%,词错误率(WER)从52.69%降至23.20%。

原文摘要 · Abstract (English)

Current movie dubbing technology can generate the desired voice from a given speech prompt, ensuring good synchronization between speech and visuals while accurately conveying the intended emotions. However, in movie dubbing, key aspects such as adapting to different dubbing styles, handling dialogue, narration, and monologue effectively, and understanding subtle details like the age and gender of speakers, have not been well studied. To address this challenge, we propose a framework of multi-modal large language model. First, it utilizes multimodal Chain-of-Thought (CoT) reasoning methods on visual inputs to understand dubbing styles and fine-grained attributes. Second, it generates high-quality dubbing through large speech generation models, guided by multimodal conditions. Additionally, we have developed a movie dubbing dataset with CoT annotations. The evaluation results demonstrate a performance improvement over state-of-the-art methods across multiple datasets. In particular, for the evaluation metrics, the SPK-SIM and EMO-SIM increases from 82.48% to 89.74%, 66.24% to 78.88% for dubbing setting 2.0 on V2C Animation dataset, LSE-D and MCD-SL decreases from 14.79 to 14.63, 5.24 to 4.74 for dubbing setting 2.0 on Grid dataset, SPK-SIM increases from 64.03 to 83.42 and WER decreases from 52.69% to 23.20% for initial reasoning setting on proposed CoT-Movie-Dubbing dataset in the comparison with the state-of-the art models.

电影配音多模态推理语音生成风格适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。