arXiv:2608.20905cs.CV2026-08

构建了中文情感对话最大规模数据集,真实还原面对面交流中的情绪表达。

EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

论文配图:EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue
图 1 · 摘自论文原文
  • 采用无干扰采集框架,让演员在20个日常场景中自然表现18类情绪。
  • 含400小时录音,情绪分布与真实人类统计偏差仅0.64,远超旧数据集。
  • 适合做多模态情感识别、语音情感分析及跨模态融合研究的团队使用。

面对面的视听交互是人类沟通的核心,传递丰富的感情与社会线索。然而,现有多模态对话数据集仍受限于情绪标注不足、情感多样性差和规模小等问题。我们提出 EmotionDialogCN,一个大规模的视听情感数据集,旨在捕捉真实的面对面交流。该数据集包含由119名专业演员在20个日常情景中完成的21,880段对话,覆盖18种情绪类别,总录制时长超过400小时,是同类中规模最大、最全面的数据集。创新的数据采集框架最大限度减少了设备干扰,确保了自然且细腻的情感表达。EmotionDialogCN 的情绪分布偏差仅为0.64(相较于真实人类统计),显著优于先前数据集的5.65;主体帧占比稳定在52%-59%之间。上述特性使该数据集在声学、词汇和视觉模态上均展现出稳定的单模态与多模态性能,融合结果进一步验证了强多模态对齐与跨模态互补性。

原文摘要 · Abstract (English)

Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue datasets remain limited by inadequate emotion annotations, poor emotional diversity, and small scale. We introduce EmotionDialogCN, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication. It contains 21,880 dialogue sessions performed by 119 professional actors across 20 everyday scenarios, covering 18 emotion categories with over 400 hours of recordings, the largest and most comprehensive dataset of its kind. A novel data collection framework minimizes equipment interference, enabling natural and nuanced emotional expressions. EmotionDialogCN achieves an emotion distribution deviation of 0.64 from real human emotion statistics (versus 5.65 for prior datasets) and consistent subject framing (52-59% frame occupancy). Together, these properties translate into stable unimodal and multimodal performance across acoustic, lexical, and visual modalities, with fusion results further underscoring strong multimodal alignment and cross-modal complementarity.

情感识别多模态中文数据集对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。