arXiv:2511.23115cs.CV2025-11

用文字描述图像情感,突破视觉模型的情感鸿沟

Analyzing Image Beyond Visual Aspect: Image Emotion Classification via Multiple-Affective Captioning

  • 通过多级对比损失提取情绪概念,生成带情感的文本描述
  • 在多个数据集上超越现有方法,准确率显著提升
  • 适合需要精准情感理解的图像分析场景

图像情感分类(IEC)是深度学习推动下的长期研究课题。尽管近年工作利用预训练视觉模型的知识,但受限于‘情感鸿沟’,其效果受限。心理学研究表明语言具有高度多样性与丰富信息,可有效弥合该鸿沟。受此启发,本文提出基于多情感描述的图像情感分类方法(ACIEC),仅依赖文本实现情感分类,充分捕捉图像中的情感信息。方法设计分层多级对比损失以检测情绪概念,并引入情感属性思维链推理生成情感语句。随后利用预训练语言模型融合情绪概念与情感句子完成分类。此外,采用基于语义相似性采样的对比损失,缓解情感数据集中类内差异大、类间差异小的问题。同时,首次考虑包含嵌入文本的图像,此前研究常忽略此类样本。大量实验表明,本方法能有效弥合情感鸿沟,在多个基准上取得更优性能。

原文摘要 · Abstract (English)

Image emotion classification (IEC) is a longstanding research field that has received increasing attention with the rapid progress of deep learning. Although recent advances have leveraged the knowledge encoded in pre-trained visual models, their effectiveness is constrained by the "affective gap" , limits the applicability of pre-training knowledge for IEC tasks. It has been demonstrated in psychology that language exhibits high variability, encompasses diverse and abundant information, and can effectively eliminate the "affective gap". Inspired by this, we propose a novel Affective Captioning for Image Emotion Classification (ACIEC) to classify image emotion based on pure texts, which effectively capture the affective information in the image. In our method, a hierarchical multi-level contrastive loss is designed for detecting emotional concepts from images, while an emotional attribute chain-of-thought reasoning is proposed to generate affective sentences. Then, a pre-trained language model is leveraged to synthesize emotional concepts and affective sentences to conduct IEC. Additionally, a contrastive loss based on semantic similarity sampling is designed to solve the problem of large intra-class differences and small inter-class differences in affective datasets. Moreover, we also take the images with embedded texts into consideration, which were ignored by previous studies. Extensive experiments illustrate that our method can effectively bridge the affective gap and achieve superior results on multiple benchmarks.

图像情感多模态语言模型对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。