对比人类与多模态大模型对视频情绪的感知,发现模型在类别层面能准确捕捉复杂情绪结构。
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
- 用情感结构匹配方法比较人类与模型对视频情绪的响应
- 模型在类别层面与人类情绪结构高度相似,单个视频匹配度较低
- 适合关注情绪理解、人机共情的AI研究者参考
近期研究揭示人类情绪具有高维、复杂的结构特征。传统模型因忽略高维性,可能遗漏情绪关键细节。本文考察了最新一代多模态大语言模型(MLLMs,如Gemini、GPT)在捕捉此类复杂情绪结构方面的能力与局限。通过对比观众观看视频后的自评情绪评分与模型生成的情绪估计,我们不仅评估了单个视频层面的表现,还分析了跨视频间情绪关系构成的整体结构。结果表明,人类与模型推断的情绪结构在整体相关性上表现出强相似性。为进一步区分是单项匹配还是粗粒度类别匹配更有效,我们采用格罗莫夫-沃瑟斯坦最优传输(Gromov Wasserstein Optimal Transport)进行分析。结果显示:尽管在严格单一项目层面表现不佳,但在引发相似情绪的视频类别层面,模型表现显著,说明其能在类别层级准确推断人类情感体验。研究提示当前最先进的MLLMs虽在单例精度上有限,但已能广泛捕捉高维情绪结构的类别级特征。
原文摘要 · Abstract (English)
Recent studies have revealed that human emotions exhibit a high-dimensional, complex structure. A full capturing of this complexity requires new approaches, as conventional models that disregard high dimensionality risk overlooking key nuances of human emotions. Here, we examined the extent to which the latest generation of rapidly evolving Multimodal Large Language Models (MLLMs) capture these high-dimensional, intricate emotion structures, including capabilities and limitations. Specifically, we compared self-reported emotion ratings from participants watching videos with model-generated estimates (e.g., Gemini or GPT). We evaluated performance not only at the individual video level but also from emotion structures that account for inter-video relationships. At the level of simple correlation between emotion structures, our results demonstrated strong similarity between human and model-inferred emotion structures. To further explore whether the similarity between humans and models is at the signle item level or the coarse-categorical level, we applied Gromov Wasserstein Optimal Transport. We found that although performance was not necessarily high at the strict, single-item level, performance across video categories that elicit similar emotions was substantial, indicating that the model could infer human emotional experiences at the category level. Our results suggest that current state-of-the-art MLLMs broadly capture the complex high-dimensional emotion structures at the category level, as well as their apparent limitations in accurately capturing entire structures at the single-item level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。