arXiv:2511.22929cs.CVcs.CL2025-11中稿 · publication at the…被引 3

测试主流视觉语言模型对艺术情绪与符号的识别能力

Artwork Interpretation with Vision Language Models: A Case Study on Emotions and Emotion Symbols

  • 用三款VLM模型分四层提问艺术作品的情感表达
  • 模型对具体图像情绪识别准确,抽象符号识别失败
  • 答案不一致问题仍存在,适合艺术分析初探者使用

情绪是艺术表达的核心。由于其抽象性,艺术中的情感表现形式多样且随历史演变,分析需艺术史专业知识。本文研究当前(2025年)视觉语言模型(VLMs)在艺术情绪表达识别上的能力。以三种VLM(Llava-Llama和两个Qwen模型)为对象,提出四类递增复杂度问题:图像一般内容、情感内容、情感表达方式、情感符号。通过专家定性评估发现,模型对图像内容识别出色,常能判断所描绘的情绪及其表达方式;但在高度抽象或象征性图像上表现不佳,符号识别仍难以可靠完成。此外,模型仍存在大语言模型常见的相关问题回答不一致现象。

原文摘要 · Abstract (English)

Emotions are a fundamental aspect of artistic expression. Due to their abstract nature, there is a broad spectrum of emotion realization in artworks. These are subject to historical change and their analysis requires expertise in art history. In this article, we investigate which aspects of emotional expression can be detected by current (2025) vision language models (VLMs). We present a case study of three VLMs (Llava-Llama and two Qwen models) in which we ask these models four sets of questions of increasing complexity about artworks (general content, emotional content, expression of emotions, and emotion symbols) and carry out a qualitative expert evaluation. We find that the VLMs recognize the content of the images surprisingly well and often also which emotions they depict and how they are expressed. The models perform best for concrete images but fail for highly abstract or highly symbolic images. Reliable recognition of symbols remains fundamentally difficult. Furthermore, the models continue to exhibit the well-known LLM weakness of providing inconsistent answers to related questions.

视觉语言模型艺术分析情绪识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。