arXiv:2511.20106cs.CL2025-11

构建多语言情感混合识别数据集,支持跨语言情感分析。

EM2LDL: A Multilingual Speech Corpus for Mixed Emotion Recognition through Label Distribution Learning

  • 构建中英粤三语混合情绪数据集,包含32类细粒度情绪分布标注。
  • 基于自监督模型的基线实验在跨性别、跨年龄、跨性格评估中表现优异。
  • 适用于心理健康监测与跨文化情感计算系统研发。

本研究提出EM2LDL,一个面向多语言混合情绪识别的新型语音语料库,旨在通过标签分布学习推动该领域发展。针对现有语料库普遍存在单语种、单标签标注、难以建模混合情绪且生态效度不足的问题,EM2LDL收录了英语、普通话和粤语的表达性语句,涵盖香港与澳门地区常见的句内语言切换现象。数据来源于在线平台的自发情绪表达,标注覆盖32个情绪类别,采用细粒度标签分布形式。基于自监督学习模型的基线实验在独立说话人条件下,实现了跨性别、跨年龄及跨人格特征的稳健性能,其中HuBERT-large-EN表现最佳。该语料库融合语言多样性与生态真实性,为多语言环境下复杂情感动态研究提供了有效工具。研究成果可应用于情感计算中的心理健康监测与跨文化交流系统。数据集、标注与基线代码已公开于https://github.com/xingfengli/EM2LDL。

原文摘要 · Abstract (English)

This study introduces EM2LDL, a novel multilingual speech corpus designed to advance mixed emotion recognition through label distribution learning. Addressing the limitations of predominantly monolingual and single-label emotion corpora \textcolor{black}{that restrict linguistic diversity, are unable to model mixed emotions, and lack ecological validity}, EM2LDL comprises expressive utterances in English, Mandarin, and Cantonese, capturing the intra-utterance code-switching prevalent in multilingual regions like Hong Kong and Macao. The corpus integrates spontaneous emotional expressions from online platforms, annotated with fine-grained emotion distributions across 32 categories. Experimental baselines using self-supervised learning models demonstrate robust performance in speaker-independent gender-, age-, and personality-based evaluations, with HuBERT-large-EN achieving optimal results. By incorporating linguistic diversity and ecological validity, EM2LDL enables the exploration of complex emotional dynamics in multilingual settings. This work provides a versatile testbed for developing adaptive, empathetic systems for applications in affective computing, including mental health monitoring and cross-cultural communication. The dataset, annotations, and baseline codes are publicly available at https://github.com/xingfengli/EM2LDL.

多语言情感识别语音数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。