让语音合成更自然,通过多模态上下文理解说话重点
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
- 融合文本与声音上下文,分层建模全局与局部语义
- 利用多模态多尺度信息提升语气强调的准确性
- 适合需要情感化语音交互的研究与产品开发
对话式文本转语音(CTTS)旨在对话场景中准确表达话语风格,日益受到关注。然而,以往研究未充分探索语音强调表达,这在人机交互中对传达意图和态度至关重要,主要受限于对话强调数据集稀缺及上下文理解困难。本文提出新型强调渲染方案ER-CTTS,包含两部分:1)同时考虑文本与声学上下文,结合全局与局部语义建模,全面理解对话背景;2)深度整合多模态与多尺度上下文,学习上下文对当前话语强调的影响。最终将推断出的强调特征输入神经语音合成器生成对话语音。为缓解数据不足问题,我们在现有对话数据集DailyTalk上添加了强调强度标注。客观与主观评估均表明,该模型在对话场景中的强调表达上优于基线模型。代码与音频样本见https://github.com/CodeStoreTTS/ER-CTTS。
原文摘要 · Abstract (English)
Conversational Text-to-Speech (CTTS) aims to accurately express an utterance with the appropriate style within a conversational setting, which attracts more attention nowadays. While recognizing the significance of the CTTS task, prior studies have not thoroughly investigated speech emphasis expression, which is essential for conveying the underlying intention and attitude in human-machine interaction scenarios, due to the scarcity of conversational emphasis datasets and the difficulty in context understanding. In this paper, we propose a novel Emphasis Rendering scheme for the CTTS model, termed ER-CTTS, that includes two main components: 1) we simultaneously take into account textual and acoustic contexts, with both global and local semantic modeling to understand the conversation context comprehensively; 2) we deeply integrate multi-modal and multi-scale context to learn the influence of context on the emphasis expression of the current utterance. Finally, the inferred emphasis feature is fed into the neural speech synthesizer to generate conversational speech. To address data scarcity, we create emphasis intensity annotations on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in emphasis rendering within a conversational setting. The code and audio samples are available at https://github.com/CodeStoreTTS/ER-CTTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。