构建多标签情感数据集,更真实评估视觉情感模型能力。
MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models

- 让20人每人从图片中选所有感受的情绪,得票聚合生成多标签数据。
- 含10344张图、236998个有效投票,覆盖8类情绪,标注更全面。
- 发现当前大模型在情绪分布预测上仍有提升空间,适合评测视觉情感模型。
本文提出一个面向多模态大语言模型(MLLMs)视觉情感分析的多标签基准数据集,用于全面评估模型对图像引发情感的预测能力。现有数据集采用单候选情绪标注方式,即每名标注者仅面对一种情绪判断是否被激发,此方法受限于图像可能引发多种情绪的事实,导致评估结果低估模型潜力。为此,本研究为每张图片雇佣20名标注者,要求其从图片中选出所有感受到的情绪,再汇总各情绪得票,形成更可靠、更具代表性的多标签数据集。最终数据集包含10,344张图像和236,998个有效投票,涵盖8种情绪。基于该数据集,我们评估了Qwen3-VL、GPT、Gemini和Claude等近期模型在主导情绪与情绪分布预测上的表现。结果表明,尽管模型取得进展,但仍存在显著改进空间。此外,实验显示以LLM为评判者的方法在主观性任务中并不总能提升性能,揭示其局限性。
原文摘要 · Abstract (English)
This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language models (MLLMs) to predict the emotions evoked by images. Recent user studies report an unintuitive finding: humans may prefer the predictions of MLLMs over the labels in existing datasets. We argue that this phenomenon stems from the suboptimal annotation scheme used in existing datasets, where each annotator is shown a single candidate emotion for each image and judges whether it is evoked or not. This approach is clearly limited because a single image can evoke multiple emotions with varying intensities. As a result, evaluations based on these datasets may underestimate the capabilities of MLLMs, yet an appropriate benchmark for evaluating such models remains lacking. To address this issue, we introduce a new multi-label benchmark dataset for visual emotion analysis toward MLLMs evaluation. We hire $20$ annotators per image and ask them to select all emotions they feel from an image. Then, we aggregate the votes across all annotators, providing a more reliable and representative dataset labeled with a distribution of emotions. The resulting dataset contains $10,344$ images with $236,998$ valid votes across eight emotions. Based on this benchmark dataset, we evaluate several recent models, including Qwen3-VL, OpenAI's GPT, Gemini, and Claude. We assess model performance on both dominant emotion prediction and emotion distribution prediction. Our results demonstrate the progress achieved by recent MLLMs while also indicating that substantial room for improvement remains. Furthermore, our experiments with LLM-as-a-judge show that the method does not consistently improve MLLMs' performance, indicating its limitations for the subjective task of visual emotion analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。