arXiv:2409.03521cs.CV2024-09被引 13

大模型能像专家一样判断画作风格、作者和年代吗?

Have Large Vision-Language Models Mastered Art History?

  • 用CLIP、LLaVA、GPT-4o三模型零样本分析画作风格与作者
  • 在两个艺术数据集上测试,准确率接近人类专家水平
  • 揭示模型对提示敏感,适合艺术史研究者参考

大型视觉语言模型(VLMs)在多个领域的图像分类中已建立新基准。本文探讨其多模态推理能力是否足以应对艺术史专家擅长的挑战:判断画作的风格、作者和创作年代。与自然图像不同,艺术品具有复杂多变的构图与风格,需上下文与风格解读,而非简单物体识别。本文首次系统评估三种VLMs——CLIP、LLaVA和GPT-4o——在零样本条件下对艺术风格、作者和时代时期的分类能力。基于两个艺术图像基准数据集,评估模型的风格理解力、提示敏感性及错误案例。同时通过对比人类艺术史专家的误判,深入分析模型的推理模式与分类特征。

原文摘要 · Abstract (English)

The emergence of large Vision-Language Models (VLMs) has established new baselines in image classification across multiple domains. We examine whether their multimodal reasoning can also address a challenge mastered by human experts. Specifically, we test whether VLMs can classify the style, author and creation date of paintings, a domain traditionally mastered by art historians. Artworks pose a unique challenge compared to natural images due to their inherently complex and diverse structures, characterized by variable compositions and styles. This requires a contextual and stylistic interpretation rather than straightforward object recognition. Art historians have long studied the unique aspects of artworks, with style prediction being a crucial component of their discipline. This paper investigates whether large VLMs, which integrate visual and textual data, can effectively reason about the historical and stylistic attributes of paintings. We present the first study of its kind, conducting an in-depth analysis of three VLMs, namely CLIP, LLaVA, and GPT-4o, evaluating their zero-shot classification of art style, author and time period. Using two image benchmarks of artworks, we assess the models' ability to interpret style, evaluate their sensitivity to prompts, and examine failure cases. Additionally, we focus on how these models compare to human art historical expertise by analyzing misclassifications, providing insights into their reasoning and classification patterns.

艺术史视觉语言模型风格识别零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。