构建了一个平衡的多模态情感数据集,支持更可靠的跨数据集泛化。
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark

- 从3100部影视作品中提取30万段语音片段,三模态同步标注七类情绪。
- 数据集严格按剧集分片,避免内容重叠,且各类情绪分布接近均衡。
- 适合研究多模态情感识别、少数情绪建模与模型泛化能力评估的人群。
理解口语对话中的情感是情感计算的关键挑战,应用于共情AI、人机交互和心理健康监测。然而现有数据集在规模、情绪分布、模态对齐和划分策略上差异较大,影响模型跨数据集泛化和少数情绪建模。我们提出SpEmoC,一个面向对话的说话片段情感数据集,包含来自3100部英语影视作品的306,544段原始音频片段。从中精选出30,000段高质量、类别均衡的片段,具有同步的视觉、音频和文本模态,通过结合预训练模型与人工验证的混合流程标注七种情绪。SpEmoC采用严格的电影与剧集级划分,防止不同数据集间内容重叠,支持更可靠的模型泛化评估。数据集保持七类情绪近似均衡分布,涵盖恐惧、厌恶等少数情绪,有利于类别平衡学习。大量实验表明,包括域内基准测试、跨数据集迁移、低数据训练、类别不平衡分析和模态迁移在内的结果均显示,均衡数据与精细划分能提升模型在其他数据集上的稳定表现。这些结果凸显了数据集设计对鲁棒且可迁移的情感识别的重要性。
原文摘要 · Abstract (English)
Understanding human emotions in spoken conversations is a key challenge in affective computing, with applications in empathetic AI, human computer interaction, and mental health monitoring. However, existing datasets vary in scale, emotion distribution, modality alignment, and data partitioning strategies, which can influence reliable cross-dataset generalization and minority-emotion modeling. We introduce SpEmoC a Speaking segment Emotion for Conversations comprising 306,544 raw clips from 3,100 English language movies and TV series. From these, 30,000 high quality, class balanced clips are curated, featuring synchronized visual, audio, and textual modalities annotated for seven emotions through a hybrid pipeline that integrates pretrained models with human validation. SpEmoC uses strict movie- and series-level splits to prevent content overlap between split sets, allowing more reliable evaluation of model generalization. The dataset also maintains a near-balanced distribution across seven emotions, including minority classes such as Fear and Disgust, which supports more balanced learning across categories. Extensive experiments, including in-domain benchmarking, cross-dataset transfer, low-data training, class-imbalance analysis, and modality transfer show that balanced data and careful splitting lead to more stable performance across emotions when models are evaluated on other datasets. These results highlight the importance of dataset design for robust and transferable multimodal emotion recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。