构建首个高质量中文非语言发声数据集,提升语音识别的情感理解能力
MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
- 基于表演式录音构建7.55小时高保真非语言发声数据
- 包含17类均衡分布的发声类别,覆盖常见情绪与意图表达
- 支持语音转写与发声分类联合建模,适合情感计算研究者使用
主流自动语音识别(ASR)系统擅长转录词汇内容,但对嵌入在语音中的非语言发声(NVs),如叹息、笑声、咳嗽等识别能力有限。这类发声传递关键情感与意图线索,对全面理解人类交流至关重要。现有进展受限于缺乏高质量、标注精细的数据集。为此,我们提出MNV-17,一个7.55小时的表演式普通话语音数据集。与依赖模型检测的多数语料不同,MNV-17凭借其表演性质,确保了非语言发声实例的高保真度与清晰性。据我们所知,MNV-17提供了最丰富的非语言发声类别,涵盖17种常见且均衡的发声类型。我们在四种主流ASR架构上对MNV-17进行了基准测试,评估其在语义转写与发声分类上的联合性能。数据集及预训练模型检查点将公开,以推动表达性语音识别研究。
原文摘要 · Abstract (English)
Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capability is important for a comprehensive understanding of human communication, as NVs convey crucial emotional and intentional cues. Progress in NV-aware ASR has been hindered by the lack of high-quality, well-annotated datasets. To address this gap, we introduce MNV-17, a 7.55-hour performative Mandarin speech dataset. Unlike most existing corpora that rely on model-based detection, MNV-17's performative nature ensures high-fidelity, clearly articulated NV instances. To the best of our knowledge, MNV-17 provides the most extensive set of nonverbal vocalization categories, comprising 17 distinct and well-balanced classes of common NVs. We benchmarked MNV-17 on four mainstream ASR architectures, evaluating their joint performance on semantic transcription and NV classification. The dataset and the pretrained model checkpoints will be made publicly available to facilitate future research in expressive ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。