用CLIP模型预测艺术风格五原则,提升自动艺术分析能力
WP-CLIP: Leveraging CLIP to Predict Wölfflin's Principles in Visual Art
- 用标注的艺术图像微调CLIP,学习预测沃尔夫林五原则
- 在生成画作和1.8万张真实画作上验证,跨风格泛化能力强
- 为艺术风格量化分析提供新工具,适合艺术计算研究者
沃尔夫林的五个风格原则为形式分析提供了结构化框架,但现有度量方法无法有效预测全部原则。图像视觉属性的计算评估需能解读色彩、构图和主题等关键元素。近年来,视觉语言模型(VLM)在评估抽象图像属性方面表现突出,成为该任务的有力候选。本文探究预训练于大规模数据的CLIP是否能理解并预测沃尔夫林原则。结果表明,其原始模型无法捕捉此类细微风格特征。为此,我们在真实艺术图像的标注数据集上对CLIP进行微调,以预测每个原则的评分。所提出的WP-CLIP模型在GAN生成画作和Pandora-18K艺术数据集上进行评估,展示了跨多样艺术风格的泛化能力。结果证明,VLM在自动化艺术分析中具有巨大潜力。
原文摘要 · Abstract (English)
Wölfflin's five principles offer a structured approach to analyzing stylistic variations for formal analysis. However, no existing metric effectively predicts all five principles in visual art. Computationally evaluating the visual aspects of a painting requires a metric that can interpret key elements such as color, composition, and thematic choices. Recent advancements in vision-language models (VLMs) have demonstrated their ability to evaluate abstract image attributes, making them promising candidates for this task. In this work, we investigate whether CLIP, pre-trained on large-scale data, can understand and predict Wölfflin's principles. Our findings indicate that it does not inherently capture such nuanced stylistic elements. To address this, we fine-tune CLIP on annotated datasets of real art images to predict a score for each principle. We evaluate our model, WP-CLIP, on GAN-generated paintings and the Pandora-18K art dataset, demonstrating its ability to generalize across diverse artistic styles. Our results highlight the potential of VLMs for automated art analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。