arXiv:2507.13155cs.LGcs.SD2025-07被引 19

公开17小时带情绪标注的非语言发声数据集,提升语音合成表现。

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech

  • 从VoxCeleb和Expresso自动提取并人工验证非语言发声
  • 10类发声+8类情绪标注,支持高质量语音合成训练
  • 可直接微调开源模型,效果媲美闭源系统

当前情感语音合成模型受限于缺乏多样化的开源非语言发声(NVs)数据集。本文提出NonverbalTTS(NVTTS),一个17小时的公开数据集,包含10类非语言发声(如笑声、咳嗽)和8种情绪类别。数据源自VoxCeleb与Expresso,经自动化检测后由人工验证。我们设计了一套完整流程,整合自动语音识别(ASR)、NV标注、情绪分类及多标注者转录融合算法。在NVTTS上微调开源文本到语音(TTS)模型,在人类评估与自动指标(包括说话人相似度和非语言发声保真度)上达到与闭源系统CosyVoice2相当的效果。通过发布NVTTS及其标注指南,我们解决了情感语音合成研究中的关键瓶颈。数据集已开放获取:https://huggingface.co/datasets/deepvk/NonverbalTTS。

原文摘要 · Abstract (English)

Current expressive speech synthesis models are constrained by the limited availability of open-source datasets containing diverse nonverbal vocalizations (NVs). In this work, we introduce NonverbalTTS (NVTTS), a 17-hour open-access dataset annotated with 10 types of NVs (e.g., laughter, coughs) and 8 emotional categories. The dataset is derived from popular sources, VoxCeleb and Expresso, using automated detection followed by human validation. We propose a comprehensive pipeline that integrates automatic speech recognition (ASR), NV tagging, emotion classification, and a fusion algorithm to merge transcriptions from multiple annotators. Fine-tuning open-source text-to-speech (TTS) models on the NVTTS dataset achieves parity with closed-source systems such as CosyVoice2, as measured by both human evaluation and automatic metrics, including speaker similarity and NV fidelity. By releasing NVTTS and its accompanying annotation guidelines, we address a key bottleneck in expressive TTS research. The dataset is available at https://huggingface.co/datasets/deepvk/NonverbalTTS.

语音合成非语言发声数据集情绪标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。