arXiv:2505.18453cs.SDcs.AI2025-05中稿 · InterSpeech被引 4

用多模态提示定制情感语音,支持文本、图像或语音作为情感输入。

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt

  • 分离语音内容、音色、情感和语调,支持多模态情感提示。
  • 在客观与主观评测中,自然度和相似度均优于现有系统。
  • 适合需要个性化情感语音生成的研究与应用开发者。

现有零样本语音合成(ZS-TTS)系统通常依赖单一提示(如参考语音或文本描述),灵活性受限。本文提出一种基于多模态提示的定制化情感零样本语音合成系统。该系统将语音解耦为内容、音色、情感和韵律四部分,允许情感提示以文本、图像或语音形式输入。为从不同提示中提取情感信息,设计了多模态提示情感编码器;引入韵律预测器以拟合韵律分布,并提出情感一致性损失以保留预测韵律中的情感信息。采用基于扩散模型的声学模型生成目标梅尔频谱图。客观与主观实验表明,本系统在自然度和相似度上均优于现有方法。演示样本见 https://mpetts-demo.github.io/mpetts_demo/。

原文摘要 · Abstract (English)

Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibility. We propose a customized emotion ZS-TTS system based on multi-modal prompt. The system disentangles speech into the content, timbre, emotion and prosody, allowing emotion prompts to be provided as text, image or speech. To extract emotion information from different prompts, we propose a multi-modal prompt emotion encoder. Additionally, we introduce an prosody predictor to fit the distribution of prosody and propose an emotion consistency loss to preserve emotion information in the predicted prosody. A diffusion-based acoustic model is employed to generate the target mel-spectrogram. Both objective and subjective experiments demonstrate that our system outperforms existing systems in terms of naturalness and similarity. The samples are available at https://mpetts-demo.github.io/mpetts_demo/.

语音合成多模态情感控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。