arXiv:2509.06074cs.CL2025-09EMNLP被引 3

通过词级交互图建模对话细粒度语义与韵律,提升语音合成自然度。

Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis

  • 构建语义与韵律双图,捕捉词级对话交互
  • 在DailyTalk上显著提升韵律表现力
  • 适合关注语音合成细节表现的研究者

对话式语音合成(CSS)旨在通过理解多模态对话历史(MDH)生成具有自然韵律的语音。现有方法虽能建模对话历史与目标语句的句级交互,却忽视了词级层面的细粒度语义与韵律知识。为此,本文提出MFCIG-CSS,一种基于多模态细粒度上下文交互图的语音合成系统。该方法构建两个专用的多模态细粒度对话交互图:语义交互图与韵律交互图,有效编码词级语义、韵律及其对后续语句的影响。融合这些交互特征后,生成的语音在韵律表达上更自然。在DailyTalk数据集上的实验表明,MFCIG-CSS优于所有基线模型,在韵律表现力方面有显著提升。代码与语音样本见https://github.com/AI-S2-Lab/MFCIG-CSS。

原文摘要 · Abstract (English)

Conversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH). The latest work predicts the accurate prosody expression of the target utterance by modeling the utterance-level interaction characteristics of MDH and the target utterance. However, MDH contains fine-grained semantic and prosody knowledge at the word level. Existing methods overlook the fine-grained semantic and prosodic interaction modeling. To address this gap, we propose MFCIG-CSS, a novel Multimodal Fine-grained Context Interaction Graph-based CSS system. Our approach constructs two specialized multimodal fine-grained dialogue interaction graphs: a semantic interaction graph and a prosody interaction graph. These two interaction graphs effectively encode interactions between word-level semantics, prosody, and their influence on subsequent utterances in MDH. The encoded interaction features are then leveraged to enhance synthesized speech with natural conversational prosody. Experiments on the DailyTalk dataset demonstrate that MFCIG-CSS outperforms all baseline models in terms of prosodic expressiveness. Code and speech samples are available at https://github.com/AI-S2-Lab/MFCIG-CSS.

语音合成多模态图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。