构建可扩展的非语言发声数据集,提升语音系统的情感表达能力。
A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
- 用统一检测模型自动识别自然语音中的非语言发声
- 创建含38,718样本的非语言发声-38K数据集,覆盖10类发声
- 适合做语音情感生成与理解的研究者使用
非语言发声(如笑声、叹气)在人类言语中对传递情绪和意图至关重要,但现有语音系统普遍忽略此类内容,严重削弱了交流的丰富性与情感智能。当前获取非语言发声的方法要么成本高且难以扩展(依赖人工标注/录制),要么不自然(基于规则合成)。为此,我们提出一种高度可扩展的自动标注框架,能够低成本、易扩展地从自然语音中标注非语言发声,兼具多样性与自然性。该框架采用统一检测模型准确识别自然语音中的非语言发声,并通过时序-语义对齐方法将其与转录文本关联。基于此框架,我们构建并发布了NonVerbalSpeech-38K数据集,包含38,718个真实场景样本,涵盖10类非语言发声,来源于真实媒体数据。实验表明,该数据集在非语言发声生成上具有更优可控性,在理解任务中表现相当可靠。
原文摘要 · Abstract (English)
Non-verbal Vocalizations (NVs), such as laughter and sighs, are vital for conveying emotion and intention in human speech, yet most existing speech systems neglect them, which severely compromises communicative richness and emotional intelligence. Existing methods for NVs acquisition are either costly and unscalable (relying on manual annotation/recording) or unnatural (relying on rule-based synthesis). To address these limitations, we propose a highly scalable automatic annotation framework to label non-verbal phenomena from natural speech, which is low-cost, easily extendable, and inherently diverse and natural. This framework leverages a unified detection model to accurately identify NVs in natural speech and integrates them with transcripts via temporal-semantic alignment method. Using this framework, we created and released \textbf{NonVerbalSpeech-38K}, a diverse, real-world dataset featuring 38,718 samples across 10 NV categories collected from in-the-wild media. Experimental results demonstrate that our dataset provides superior controllability for NVs generation and achieves comparable performance for NVs understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。