arXiv:2605.25504eess.AS2026-05

让语音合成更真实:精细控制非语言发声表达情感

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control

论文配图:Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control
图 1 · 摘自论文原文
  • 用标签标注非语言发声的类型、频率和时长,实现精准控制
  • 情绪识别准确率达78.8%,高唤醒情绪识别超82%
  • 适合做情感化语音合成与人机交互研究者参考

当前情感文本转语音(TTS)模型虽能控制言语韵律,却常忽略对人类情感至关重要的非语言发声(NVs)。尽管近期出现了一些非语言数据集,但普遍缺乏高质量、细粒度的标注,限制了模型对非语言发声生成的精确控制。为此,我们提出一种新的细粒度非语言表达合成方法:从EARS语料库中重新整理并处理女性非语言发声样本,设计基于标签的新标注方案以编码非语言发声的类型、频率与持续时间,并构建一个情感TTS基准测试以验证其有效性。评估显示,尽管我们的方法在感知自然度上略有下降,但显著提升了表达力(eMOS 4.20)和情绪识别准确率(78.8%)。情感特异性分析表明,非语言线索对高唤醒情绪(如快乐82.5%、恐惧82.7%)极为有效,几乎完美传达悲伤(98.3%)。

原文摘要 · Abstract (English)

While current emotional Text-to-Speech (TTS) models have successfully controlled verbal prosody, they often ignore non-verbal vocalizations (NVs), which are essential for authentic human emotion. Although some non-verbal datasets have recently emerged, they often lack high-quality, fine-grained annotations, which restricts a model's ability to precisely control NV generation. To address this limitation, we propose a novel approach for fine-grained non-verbal expression synthesis. We curate and reprocess female NV utterances from the EARS corpus, develop a new annotation scheme using tags to encode NV types, frequencies, and durations, and build an emotional TTS benchmark to demonstrate its effectiveness. Our evaluation shows that while our NV approach leads to minor trade-offs in perceived naturalness, it significantly improves expressiveness (eMOS 4.20) and emotional recognition accuracy (78.8%). Emotion-specific analysis further reveals that NV cues are highly effective for high-arousal emotions like happy (82.5%) and fear (82.7%), and almost perfectly convey sadness (98.3%).

语音合成情感表达非语言发声TTS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。