构建首个细粒度语音情感识别基准,支持40类情绪与敏感情感研究。
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
- 用合成语音实现40类情绪数据采集,兼顾隐私与伦理。
- 4700样本经三名专家验证,准确率达95%(愤怒)。
- 适合做精细情感识别、人机共情系统的研究者。
语音情感识别(SER)受限于现有数据集通常仅覆盖6-10种基本情绪,规模小、多样性不足,且收集敏感情绪存在伦理问题。我们提出EMONET-VOICE,包含两个部分:(1) EmoNet-Voice Big,一个5000小时的多语言预训练数据集,涵盖40种细粒度情绪类别,来自11位语音人和4种语言;(2) EmoNet-Voice Bench,一个4700样本的严格验证基准,所有情绪存在性与强度水平均获心理学专家一致共识。通过最先进的合成语音生成技术,我们的隐私保护方法实现了对敏感情绪(如疼痛、羞耻)的伦理化纳入,并保持可控实验条件。每条样本经三位心理学专家验证。实验表明,基于该数据训练的共情模型在EmoDB和RAVDESS上表现良好,具备强泛化能力。综合评估显示,高唤醒情绪(如愤怒)识别准确率达95%,但感知相似的情绪(如悲伤与痛苦)区分率仅为63%,为推进细腻情感人工智能提供可量化指标。EMONET-VOICE确立了大规模、伦理合规、细粒度语音情感研究的新范式。
原文摘要 · Abstract (English)
Speech emotion recognition (SER) systems are constrained by existing datasets that typically cover only 6-10 basic emotions, lack scale and diversity, and face ethical challenges when collecting sensitive emotional states. We introduce EMONET-VOICE, a comprehensive resource addressing these limitations through two components: (1) EmoNet-Voice Big, a 5,000-hour multilingual pre-training dataset spanning 40 fine-grained emotion categories across 11 voices and 4 languages, and (2) EmoNet-Voice Bench, a rigorously validated benchmark of 4,7k samples with unanimous expert consensus on emotion presence and intensity levels. Using state-of-the-art synthetic voice generation, our privacy-preserving approach enables ethical inclusion of sensitive emotions (e.g., pain, shame) while maintaining controlled experimental conditions. Each sample underwent validation by three psychology experts. We demonstrate that our Empathic Insight models trained on our synthetic data achieve strong real-world dataset generalization, as tested on EmoDB and RAVDESS. Furthermore, our comprehensive evaluation reveals that while high-arousal emotions (e.g., anger: 95% accuracy) are readily detected, the benchmark successfully exposes the difficulty of distinguishing perceptually similar emotions (e.g., sadness vs. distress: 63% discrimination), providing quantifiable metrics for advancing nuanced emotion AI. EMONET-VOICE establishes a new paradigm for large-scale, ethically-sourced, fine-grained SER research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。