首个系统评估视觉语言模型识图情绪能力,揭示其优劣与错误原因。
Evaluating Vision-Language Models for Emotion Recognition
- 构建图像诱发情绪识别基准,全面测试模型表现。
- 发现模型准确率受场景语义与情感强度影响显著。
- 通过人工评估定位错误根源,指导未来情感研究方向。
大型视觉语言模型(VLMs)在多项客观多模态推理任务中取得突破性进展。然而,为提升其与人类共情及有效沟通的能力,增强对情感的理解至关重要。尽管情感理解受到广泛关注,但针对该任务的细致评估仍显不足,难以支撑下游微调工作。本文首次系统评估VLMs在图像诱发情绪识别任务中的表现,构建了相关基准,并从准确性与鲁棒性角度分析模型性能。通过多项实验,揭示了情绪识别表现依赖的关键因素,刻画了模型在识别过程中产生的各类错误。最后,结合人工评估研究,明确错误潜在成因。基于实验结果,提出未来在VLM情感研究方面的改进建议。
原文摘要 · Abstract (English)
Large Vision-Language Models (VLMs) have achieved unprecedented success in several objective multimodal reasoning tasks. However, to further enhance their capabilities of empathetic and effective communication with humans, improving how VLMs process and understand emotions is crucial. Despite significant research attention on improving affective understanding, there is a lack of detailed evaluations of VLMs for emotion-related tasks, which can potentially help inform downstream fine-tuning efforts. In this work, we present the first comprehensive evaluation of VLMs for recognizing evoked emotions from images. We create a benchmark for the task of evoked emotion recognition and study the performance of VLMs for this task, from perspectives of correctness and robustness. Through several experiments, we demonstrate important factors that emotion recognition performance depends on, and also characterize the various errors made by VLMs in the process. Finally, we pinpoint potential causes for errors through a human evaluation study. We use our experimental results to inform recommendations for the future of emotion research in the context of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。