用多模态大模型生成图像描述再分析情感,准确率比传统方法高64.8%。
Multimodal LLMs See Sentiment
- 先让多模态大模型描述图像,再用语言模型分析情感
- 在多个数据集上准确率提升最多达64.8%(相比CNN基线)
- 无需目标数据训练即可跨数据集表现优于现有方法
在以图像为主导的数字环境中,理解视觉内容所传达的情感日益重要。然而,情感感知依赖复杂的场景级语义,对计算模型构成挑战。本文通过系统性评估,从三方面研究多模态大语言模型(MLLMs)在图像情感分析中的表现:(i) 直接使用MLLM进行图像情感分类;(ii) 利用预训练语言模型分析MLLM生成的描述;(iii) 在标注情感的描述上微调语言模型以评估性能与泛化能力。实验表明,两阶段的描述中介管道在多种设置下显著提升准确率,尤其当语言模型被微调时效果更佳。在不同一致度阈值和情感粒度下,该方案在基准测试中分别超越词典法、CNN和Transformer基线高达30.9%、64.8%和42.4%。跨数据集评估中,无需在目标数据集上训练或微调,仍超过最佳域内基线超8%。研究全面评估了描述中介情感分析的有效性与局限性,并提供可复现的基准资源。
原文摘要 · Abstract (English)
Understanding how visual content conveys sentiment is increasingly important in a digital landscape dominated by imagery. However, sentiment perception depends on complex scene-level semantics, making this a challenging task for computational models. This paper examines how Multimodal Large Language Models (MLLMs) perform sentiment analysis in images through a systematic, evaluation-driven study encompassing three perspectives: (i) direct sentiment classification from images using MLLMs; (ii) sentiment analysis on MLLM-generated descriptions using pre-trained LLMs; and (iii) fine-tuning these LLMs on sentiment-labeled descriptions to assess performance and generalization. Experiments on a recent benchmark show that a two-stage MLLM description-mediated pipeline can substantially improve prediction accuracy under several evaluation settings, particularly when the LLM component is fine-tuned. Across different agreement thresholds and sentiment granularities, the strongest configurations of this pipeline outperform lexicon-, CNN-, and Transformer-based baselines in our benchmark by up to 30.9%, 64.8%, and 42.4%, respectively. In cross-dataset evaluation, the proposed pipeline - without training or fine-tuning on the target dataset - still surpasses the best in-domain baseline by over 8%. Overall, the study provides a comprehensive assessment of MLLM description-mediated sentiment analysis, clarifying the conditions under which it is effective, the scenarios in which it fails, and its comparison with traditional vision-based approaches, while also providing a reproducible benchmark resource for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。