arXiv:2607.19836cs.IRcs.DL2026-07

用词库结构分析CLIP在历史照片检索中的失败原因。

Using Hierarchical Controlled Vocabularies to Understand CLIP Retrieval Failures in Historical Photo Collections

  • 基于词库的层级结构,分析检索失败与词项分类和深度的关系。
  • 发现视觉聚类与图文对齐几乎无关,共同揭示不同失败模式。
  • 微调提升整体性能,尤其对浅层词项效果更显著。

GLAM机构使用如艺术与建筑主题词表(AAT)等受控词库组织图像访问。尽管像CLIP这样的视觉-语言模型被广泛用于内容驱动的图像检索,其性能却存在差异。已有研究指出该差异与概念抽象程度和出现频率相关,但尚未从词库自身结构属性角度解释。AAT将概念分为对象、活动、人物等大类,并在类内进行层级排列。本文探究两个结构属性(根类类型与层级深度)是否能解释CLIP检索的成功与失败,以及微调的有效性。在三个标注了AAT术语的历史摄影集合上,我们考察了视觉一致性(术语对应的图片在CLIP嵌入空间中是否聚集)、图文对齐(标签是否靠近该聚集区)及标准检索指标(二者混合)。结果表明,视觉一致性和图文对齐几乎不相关,且联合区分出不同的失败模式。那些图片高度聚集但标签远离聚集区的词项,在所有数据集中检索表现差,其中两个集合中甚至比双失败词项更差。虽然标准检索指标与结构属性无显著相关性,但根类类型显著区分了具有不同视觉一致性的类别。此外,微调总体改善了检索表现,但收益主要集中在层级较浅的词项上,其图文对齐提升幅度超过概念频率所能解释的部分。

原文摘要 · Abstract (English)

GLAM institutions (Galleries, Libraries, Archives, and Museums) organise image access using controlled vocabularies such as the Art and Architecture Thesaurus (AAT). For content-based image retrieval in these settings, vision-language models like CLIP are increasingly used, but their performance varies. This variation is known to relate to measures like concept abstraction and concept frequency. However, no prior work explains this variation in terms of the structural properties of vocabularies like the AAT that GLAM professionals already use. The AAT groups concepts into broad facets (Objects, Activities, Agents, etc.) and arranges terms hierarchically within them. In this paper, we ask whether two structural properties (root facet type and hierarchy depth) explain where CLIP retrieval succeeds and fails, and where fine-tuning helps. Across three historical photographic collections annotated with AAT terms, we examine visual coherence (whether a term's photographs cluster in CLIP's embedding space), text-image alignment (whether its label is near that cluster), and standard retrieval measures, which conflate the two. We find that visual coherence and text-image alignment are nearly uncorrelated across terms and jointly separate distinct failure modes. Terms whose photographs cluster tightly but whose label is distant from the cluster retrieve poorly in every collection, in two of three collections even worse than terms that fail on both metrics. We also show that while retrieval metrics do not correlate significantly with either structural property, root facet type does significantly separate categories with varying visual coherence. Finally, we find that fine-tuning improves retrieval overall, but its gains favour shallower terms in the hierarchy, where text-image alignment improves most, beyond what concept frequency explains.

CLIP图像检索词库结构历史图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。