arXiv:2604.23348cs.CVcs.AI2026-04

构建多模态情感动态理解基准,评估模型对情绪变化的推理与预测能力。

EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs

论文配图:EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs
图 1 · 摘自论文原文
  • 设计四类任务,从情绪变化检测到未来情绪预测,形成渐进式评估框架。
  • 包含1000段视频、3000+问答对,覆盖12种真实社交场景。
  • 揭示当前模型在多人复杂情境下推理能力不足,尤其缺乏精细情绪建模。

近年来,多模态大语言模型(MLLMs)在感知、推理和生成方面展现出强大能力,日益应用于社交机器人与人机交互等场景,而理解人类情绪至关重要。然而现有基准大多将情绪理解视为静态识别问题,未能考察模型是否能把握情绪作为动态演变过程的能力——即情绪如何转变、在不同社会情境中展开。为填补这一空白,我们提出EmoTrans,一个用于评估多模态视频中情绪动态理解的基准。EmoTrans包含1,000段精心采集并人工标注的视频片段,涵盖12种真实场景,并提供超过3,000个特定任务的问答对,支持细粒度评估。该基准引入四项任务:情绪变化检测(ECD)、情绪状态识别(ESI)、情绪过渡推理(ETR)和下一情绪预测(NEP),构成从粗粒度检测到深层推理与预测的渐进评估体系。我们对18个先进MLLMs进行了全面评估,发现尽管当前模型在粗粒度情绪变化检测上表现尚可,但在精细情绪动态建模方面仍存在明显短板;尤其在多人复杂情境中,性能显著下降,且以推理为导向的模型并未带来稳定提升。为推动后续研究,我们已公开发布基准数据、评估协议与代码,地址为https://github.com/Emo-gml/EmoTrans。

原文摘要 · Abstract (English)

Recent multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and generation, and are increasingly used in applications such as social robots and human-computer interaction, where understanding human emotions is essential. However, existing benchmarks mainly formulate emotion understanding as a static recognition problem, leaving it largely unclear whether current MLLMs can understand emotion as a dynamic process that evolves, shifts between states, and unfolds across diverse social contexts. To bridge this gap, we present EmoTrans, a benchmark for evaluating emotion dynamics understanding in multimodal videos. EmoTrans contains 1,000 carefully collected and manually annotated video clips, covering 12 real-world scenarios, and further provides over 3,000 task-specific question-answer (QA) pairs for fine-grained evaluation. The benchmark introduces four tasks, namely Emotion Change Detection (ECD), Emotion State Identification (ESI), Emotion Transition Reasoning (ETR), and Next Emotion Prediction (NEP), forming a progressive evaluation framework from coarse-grained detection to deeper reasoning and prediction. We conduct a comprehensive evaluation of 18 state-of-the-art MLLMs on EmoTrans and obtain two main findings. First, although current MLLMs show relatively stronger performance on coarse-grained emotion change detection, they still struggle with fine-grained emotion dynamics modeling. Second, socially complex settings, especially multi-person scenarios, remain substantially challenging, while reasoning-oriented variants do not consistently yield clear improvements. To facilitate future research, we publicly release the benchmark, evaluation protocol, and code at https://github.com/Emo-gml/EmoTrans.

多模态情绪理解动态建模基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。