让语音合成更真实:自动生成符合情绪的笑声、叹气等非语言声音
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
- 基于小规模数据集,通过增强策略扩展非语言声音类型和插入位置
- 生成的非语言声音更丰富多样,且与语义情感高度匹配
- 适合需要高情感表现力的语音合成应用,如虚拟助手、游戏角色
非语言声音(NVs),如笑声和叹气,在情感语音合成中对表达情绪至关重要。然而,在开放场景下,由于非语言声音数据有限且缺乏显式标注,学习多样化且上下文对齐的非语言声音仍具挑战。为此,我们提出 Affectron 框架,实现情感一致且上下文对齐的非语言声音生成。该框架基于一个小型开源解耦语料库,引入非语言声音增强训练策略,拓展了非语言声音类型和插入位置的分布。进一步将非语言声音结构掩码机制融入仅以言语内容预训练的语音主干网络,实现多样化且自然的非语言声音合成。实验表明,Affectron 生成的非语言声音更具表现力和多样性,同时保持了言语流的自然性。
原文摘要 · Abstract (English)
Nonverbal vocalizations (NVs), such as laughter and sighs, are central to the expression of affective cues in emotional speech synthesis. However, learning diverse and contextually aligned NVs remains challenging in open settings due to limited NV data and the lack of explicit supervision. Motivated by this challenge, we propose Affectron as a framework for affective and contextually aligned NV generation. Built on a small-scale open and decoupled corpus, Affectron introduces an NV-augmented training strategy that expands the distribution of NV types and insertion locations. We further incorporate NV structural masking into a speech backbone pre-trained on purely verbal speech to enable diverse and natural NV synthesis. Experimental results demonstrate that Affectron produces more expressive and diverse NVs than baseline systems while preserving the naturalness of the verbal speech stream.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。