构建可扩展的多模态情绪理解基准,区分表达与唤起情绪。
E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

- 基于贝叶斯成对对齐,用少量标注生成连续情绪估计。
- 包含2524个视频、12314个问答对,覆盖三类情绪评估任务。
- 适合研究多模态情感智能的学者,尤其关注细粒度情绪识别。
理解表达与唤起情绪对多模态大语言模型实现全面情感感知交互至关重要。现有基准通常孤立评估表达或唤起情绪,且情感刻画粗糙不全。为此,我们提出E$^3$mo-Bench,一个包含2524个视频、12314个问答对的可扩展基准,涵盖预设情感视角。通过三项互补任务:情绪感知、开放词汇识别和效价-唤醒-主导力(VAD)评估,综合评测情绪理解能力。为高效实现可靠连续标注,我们提出贝叶斯成对对齐方法,将稀疏、低负担的成对判断聚合为锚定参考的VAD估计。此外,我们开发了E$^3$mo-Score,一种无需训练的代理模型,通过五模型委员会聚合互补判断以提升VAD估计性能。大量实验验证了框架有效性,并揭示了唤起与表达情绪范式间显著性能差异。这些发现结合多模态大模型在细粒度识别与维度评估上的持续短板,为推进多模态情感智能指明方向。
原文摘要 · Abstract (English)
Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations. To bridge this gap, we introduce E$^3$mo-Bench, a scalable benchmark comprising $12{,}314$ question-answer pairs across $2{,}524$ videos with predefined affective perspectives. It evaluates evoked and expressed emotion understanding via $3$ complementary tasks: emotion perception, open-vocabulary recognition, and valence-arousal-dominance (VAD) assessment. To efficiently scale reliable continuous annotations, we propose Bayesian Pairwise Alignment, which aggregates sparse, low-burden pairwise judgments into anchor-referenced VAD estimates. Furthermore, we develop E$^3$mo-Score, a training-free agent that aggregates complementary judgments from a five-model committee to improve VAD estimation. Extensive experiments validate the effectiveness of our framework and expose a pronounced performance skew between evoked and expressed emotion paradigms. These findings, coupled with MLLMs' persistent deficits in fine-grained recognition and dimensional assessment, chart a clear course for advancing multimodal emotional intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。