用视觉语言模型预测图片引发的情绪,发现效果不错但有偏差。
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
- 用九种视觉语言模型零样本预测情绪类别和评分
- 六类情绪分类准确率60%-80%,12类任务60%-75%
- 对愤怒和惊讶预测最差,且高估情绪强度
视觉语言模型(VLMs)在大规模推断视觉刺激引发的情绪方面展现出潜力,但其输出与人类情感评分的对齐程度尚不明确。我们基准测试了九种VLMs,涵盖顶尖专有模型与开源模型,在三个心理测量学验证的情感图像数据集上进行评估:国际情感图片系统(IAPS)、Nencki情感图片系统(NAPS)以及人工智能生成情感图像图书馆(LAI-GAI)。模型在零样本设置下执行两项任务:(i) 最强情绪分类(选择图像引发的最强离散情绪);(ii) 在1-7/9级李克特量表上连续预测人类对离散情绪类别及情感维度的评分。还使用去标识化参与者元数据评估了受试者条件提示对LAI-GAI数据集的影响。结果显示,离散情绪分类表现良好,六类标签准确率通常为60%至80%,十二类任务为60%至75%。所有数据集中,对愤怒和惊讶的预测准确率最低。连续评分预测与人类评分中等至强相关(r > 0.75),但存在系统性偏差,尤其在唤醒度维度表现较弱,并倾向于高估反应强度。受试者条件提示带来的预测变化微小且不一致。总体而言,VLMs能捕捉广泛情感趋势,但缺乏心理测量验证评分的细微差别,凸显其在情感计算与心理健康应用中的潜力与当前局限。
原文摘要 · Abstract (English)
Vision-language models (VLMs) show promise as tools for inferring affect from visual stimuli at scale; it is not yet clear how closely their outputs align with human affective ratings. We benchmarked nine VLMs, ranging from state-of-the-art proprietary models to open-source models, on three psycho-metrically validated affective image datasets: the International Affective Picture System, the Nencki Affective Picture System, and the Library of AI-Generated Affective Images. The models performed two tasks in the zero-shot setting: (i) top-emotion classification (selecting the strongest discrete emotion elicited by an image) and (ii) continuous prediction of human ratings on 1-7/9 Likert scales for discrete emotion categories and affective dimensions. We also evaluated the impact of rater-conditioned prompting on the LAI-GAI dataset using de-identified participant metadata. The results show good performance in discrete emotion classification, with accuracies typically ranging from 60% to 80% on six-emotion labels and from 60% to 75% on a more challenging 12-category task. The predictions of anger and surprise had the lowest accuracy in all datasets. For continuous rating prediction, models showed moderate to strong alignment with humans (r > 0.75) but also exhibited consistent biases, notably weaker performance on arousal, and a tendency to overestimate response strength. Rater-conditioned prompting resulted in only small, inconsistent changes in predictions. Overall, VLMs capture broad affective trends but lack the nuance found in validated psychological ratings, highlighting their potential and current limitations for affective computing and mental health-related applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。