arXiv:2608.13974cs.CV2026-08

通过分步聚焦视觉线索,让AI更懂艺术背后的微妙情感。

ProFocus: Interpreting Affective Experience in Artistic Images with Progressive Visual Focusing

论文配图:ProFocus: Interpreting Affective Experience in Artistic Images with Progressive Visual Focusing
图 1 · 摘自论文原文
  • 用三层语义提示引导模型逐步关注艺术图像中的氛围、叙事和细节。
  • 在ArtEmis数据集上情绪识别准确率超越现有方法,解释更贴近人类感受。
  • 适合研究艺术理解、可解释AI或人机情感交互的学者与开发者。

理解图像引发的情感反应是实现情感智能的核心。与自然图像不同,视觉艺术有意通过抽象概念和视觉隐喻激发观者情感,使情感解读尤为困难。然而,现有方法多依赖通用视觉嵌入(如CLIP),难以捕捉艺术情感背后的细微线索。为此,我们提出 extbf{ProFocus},一种基于渐进式视觉聚焦的框架,模拟人类审美认知的层级理论。技术上包含两个核心组件:层次化艺术评论器(HAC)与渐进提示融合(PHF)。HAC利用多模态大语言模型在三个认知层级——氛围风格、叙事主体、具体细节——生成结构化语言先验,将艺术感知转化为连贯语义引导。在此基础上,PHF摒弃传统跨模态融合方式,分步注入层级提示至视觉特征中,实现类人感知的渐进聚焦过程,从而捕捉微妙情感线索并生成更忠实的解释。在ArtEmis v1.0与v2.0数据集上的大量实验表明,ProFocus在情绪识别与情感解释任务中持续优于当前最佳方法。

原文摘要 · Abstract (English)

Interpreting the emotional responses triggered by images is central to achieving emotional intelligence. Compared with natural images, visual art is intentionally created to elicit emotional responses from its viewers through abstract concepts and visual metaphors, making affective interpretation particularly challenging. However, most existing methods rely on general-purpose visual embeddings (e.g., CLIP), failing to capture the nuanced cues underlying artistic emotion. To address this gap, we propose \textbf{ProFocus}, a novel framework that models affective experience in artistic images via progressive visual focusing. The key idea is to model visual representation learning inspired by a hierarchical cognitive theory of human aesthetic appreciation. Technically, ProFocus contains two core components: a Hierarchical Art Critic (HAC) and a Progressive Hint Fusion (PHF) module. HAC leverages multimodal large language models to generate structured linguistic priors at three cognitive levels--atmospheric style, narrative subjects, and concrete details--thereby translating artistic perception into coherent semantic guidance. Building upon these priors, PHF departs from conventional cross-modal fusion by sequentially injecting the hierarchical hints into visual features, enabling a progressive focusing process that mirrors human perception. This design allows the model to capture subtle affective cues and produce more faithful explanations. Extensive experiments on the ArtEmis v1.0 and v2.0 datasets demonstrate that ProFocus consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation. Project page: https://github.com/Zhang-Zhiyan/ProFocus.

艺术理解情感识别可解释AI多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。