用视觉语言模型分析艺术作品年代,揭示数据偏见对时间判断的影响
Uncertainty-Aware Art-Historical Dating with Vision-Language Models
- 将艺术品年代预测建模为带不确定性估计的回归任务
- 视觉语言模型在时间预测上优于纯视觉自监督模型
- 发现模型时间知识受收藏、数字化等机构条件影响
博物馆与档案数据集并不反映真实的艺术创作历史,而是承载了收藏、保存、编目和数字化等偶然历史。这直接影响预训练图像表示的解读:它们看似编码了历史时间,实则反映的是对象成为数据时的制度性条件。我们称此现象为时间纠缠,并通过在冻结图像嵌入上构建不确定性感知的回归任务来研究它。我们在一个时间可控的Wikidata艺术作品语料库上评估多个预训练视觉模型。结果表明,这些模型包含可用的时间信息,其中视觉-语言模型(VLMs)表现优于纯视觉自监督基线。然而定性分析显示,这种时间知识受到多种偏见的影响。
原文摘要 · Abstract (English)
Museum and archival datasets do not mirror historical artistic production, but materialize the contingent histories of collecting, preservation, cataloging, and digitization. This has direct consequences for interpreting pretrained image representations: they may appear to encode historical time while actually encoding the institutional conditions under which objects become visible as data. We describe this phenomenon as temporal entanglement and investigate it by formulating artwork dating as an uncertainty-aware regression task over frozen image embeddings. We evaluate several pretrained vision models on a temporally controlled Wikidata corpus of artworks. Our results show that these models contain usable temporal information, with Vision-Language Models (VLMs) outperforming purely visual self-supervised baselines. However, a qualitative analysis indicates that this temporal knowledge is shaped by various biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。