检验视觉模型能否真正识别艺术风格而非认出画家。
Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings
- 用画家不重叠的评估方式,避免模型靠认画家作弊。
- 风格分类准确率从0.87降至0.77,超现实主义跌20点最明显。
- 视觉结构差异揭示:某些风格更依赖共性形式而非个人笔迹。
来自CLIP等冻结图像嵌入模型常被用于艺术史风格分类,报告准确率很高。我们质疑这种高准确率是源于对风格的真实理解,还是仅靠识别特定画家。标准评估采用随机划分,导致同一画家的作品可能同时出现在训练集和测试集中,使分类器可通过认出画家而非风格得分。本文改用画家不重叠的评估协议,每次将所有某位画家作品整体排除,确保无作品基于其同画家作品被分类。在包含4个二十世纪艺术流派、共320幅画作的平衡数据集上,5-最近邻风格分类准确率从0.87降至0.77,且下降幅度极不均衡:印象派与立体主义变化微小,而超现实主义下降20个百分点。该现象在四种图像编码器(包括纯视觉自监督模型)中一致出现,说明问题根植于视觉结构而非语言。当编码器捕捉到风格共有的形式特征时,个体画家几乎无法识别,但风格仍能稳健区分;而超现实主义则表现出相反模式。因此我们主张,画家不重叠评估是衡量冻结嵌入中真实风格理解的必要条件。
原文摘要 · Abstract (English)
Frozen image embeddings from models such as CLIP are increasingly used to classify paintings by art-historical style, with high reported accuracy. We ask whether this accuracy reflects an understanding of style or the recognition of individual artists. Standard evaluation uses random splits in which works by the same artist appear on both sides, so a classifier can succeed by recognising the painter rather than the movement. We re-evaluate style classification under an artist-disjoint protocol, holding out every artist in turn so that no work is ever classified using other works by its own painter. On a balanced dataset of 320 paintings across four twentieth-century movements, 5-NN style accuracy falls from 0.87 to 0.77 under this protocol, and the drop is sharply uneven. Impressionism and Cubism barely move, while Surrealism falls twenty points. The pattern holds across four image encoders, including a vision-only self-supervised model, which places the effect in visual structure rather than language. Where an encoder captures genuine shared form, individual artists are barely recognisable yet style is robust, while Surrealism shows the opposite. We argue that artist-disjoint evaluation is necessary to measure stylistic understanding in frozen embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。