评测图文模型与多模态大模型的情绪理解能力,发现生成正向情绪更优,但整体仍不及人类。
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
- 构建7145张表情图像库,统一评估文本生成与多模态模型的情绪表现
- 现有文本生成模型更擅长生成正向情绪,且效果优于负向情绪
- 多模态模型在细粒度情绪识别上仍远未达到人类水平,适合研究情绪感知的学者
人工智能在人机交互、具身智能及虚拟数字人设计中日益关注情绪理解与表达能力。为此,本文提出MEMO-Bench,一个包含7,145张由12个文本到图像模型生成的面孔图像的综合性基准,每张图对应六种不同情绪之一。该基准首次同时支持对文本生成模型和多模态大语言模型(MLLMs)在情感分析中的评估。采用从粗粒度到细粒度的渐进式评测方法,提供更全面的分析维度。实验表明,现有文本生成模型在生成正向情绪方面表现更佳,而多模态大模型虽具备一定情绪识别能力,但在细粒度分析中仍显著低于人类水平。该基准将公开以推动相关研究。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) has demonstrated significant capabilities in various fields, and in areas such as human-computer interaction (HCI), embodied intelligence, and the design and animation of virtual digital humans, both practitioners and users are increasingly concerned with AI's ability to understand and express emotion. Consequently, the question of whether AI can accurately interpret human emotions remains a critical challenge. To date, two primary classes of AI models have been involved in human emotion analysis: generative models and Multimodal Large Language Models (MLLMs). To assess the emotional capabilities of these two classes of models, this study introduces MEMO-Bench, a comprehensive benchmark consisting of 7,145 portraits, each depicting one of six different emotions, generated by 12 Text-to-Image (T2I) models. Unlike previous works, MEMO-Bench provides a framework for evaluating both T2I models and MLLMs in the context of sentiment analysis. Additionally, a progressive evaluation approach is employed, moving from coarse-grained to fine-grained metrics, to offer a more detailed and comprehensive assessment of the sentiment analysis capabilities of MLLMs. The experimental results demonstrate that existing T2I models are more effective at generating positive emotions than negative ones. Meanwhile, although MLLMs show a certain degree of effectiveness in distinguishing and recognizing human emotions, they fall short of human-level accuracy, particularly in fine-grained emotion analysis. The MEMO-Bench will be made publicly available to support further research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。