arXiv:2412.18748cs.MMcs.CL2024-12中稿 · ICASSP 2025被引 9

通过多尺度多模态交互提升视频配音的语调表现力。

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

  • 设计双共享编码器,融合全局与局部多模态上下文特征。
  • 在Chem数据集上显著提升配音情感表达效果。
  • 适合语音合成与影视字幕生成研究者参考。

自动视频配音(AVD)旨在根据文本脚本生成与口型和面部情绪对齐的语音。现有研究虽关注多模态上下文以增强语调表现力,但忽略了两个关键问题:1)上下文中的多尺度语调属性会影响当前句子的语调;2)上下文中的语调线索会与当前句子交互,影响最终的语调表现力。为此,我们提出M2CI-Dubber,一种用于AVD的多尺度多模态上下文交互方案。该方案包含两个共享的M2CI编码器,用于建模多尺度多模态上下文,并促进其与当前句子的深度交互。通过为每种模态提取全局和局部特征,利用基于注意力的聚合与交互机制,并采用基于交互的图注意力网络进行融合,所提方法增强了当前句子合成语音的语调表现力。在Chem数据集上的实验表明,该模型在配音表现力方面优于基线方法。代码与演示可访问:https://github.com/AI-S2-Lab/M2CI-Dubber。

原文摘要 · Abstract (English)

Automatic Video Dubbing (AVD) generates speech aligned with lip motion and facial emotion from scripts. Recent research focuses on modeling multimodal context to enhance prosody expressiveness but overlooks two key issues: 1) Multiscale prosody expression attributes in the context influence the current sentence's prosody. 2) Prosody cues in context interact with the current sentence, impacting the final prosody expressiveness. To tackle these challenges, we propose M2CI-Dubber, a Multiscale Multimodal Context Interaction scheme for AVD. This scheme includes two shared M2CI encoders to model the multiscale multimodal context and facilitate its deep interaction with the current sentence. By extracting global and local features for each modality in the context, utilizing attention-based mechanisms for aggregation and interaction, and employing an interaction-based graph attention network for fusion, the proposed approach enhances the prosody expressiveness of synthesized speech for the current sentence. Experiments on the Chem dataset show our model outperforms baselines in dubbing expressiveness. The code and demos are available at \textcolor[rgb]{0.93,0.0,0.47}{https://github.com/AI-S2-Lab/M2CI-Dubber}.

视频配音多模态语调生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。