arXiv:2604.24469cs.IRcs.CV2026-04

自监督视觉表征的几何结构影响图像检索效果,越均匀越准。

Geometric Analysis of Self-Supervised Vision Representations for Semantic Image Retrieval

  • 分析自监督模型生成表征的几何特性对检索的影响
  • 高各向异性表征使近似最近邻搜索性能下降,即使线性探测准确率正常
  • 表征越均匀、局部纯净,越适合基于距离的检索索引,提升语义检索

基于内容的图像检索(CBIR)系统允许用户根据视觉内容而非元数据搜索图像。文本领域已受益于BERT等无监督方法生成的向量检索。然而,现代自监督视觉学习方法在CBIR文献中很少被报道,多数仍依赖有监督模型或多模态对齐方法。本文评估了现代自监督视觉表征在典型检索架构(使用向量数据库与最近邻搜索)下的表现。结果表明,潜在空间的几何结构会影响近似最近邻(ANN)索引性能。具体而言,若干现代自监督方法生成的高各向异性、高偏度表征会降低基于划分和哈希的搜索效率,即使其线性探测或K-NN准确率未受影响。相比之下,具有更高各向同性与局部纯净性的表征更符合ANN索引的距离假设,从而显著提升语义检索性能。

原文摘要 · Abstract (English)

Content-based image retrieval (CBIR) systems enable users to search images based on visual content instead of relying on metadata. The text domain has benefited from vector search of representations created with unsupervised methods such as BERT. However, modern self-supervised learning methods for vision are mostly not reported in CBIR-related literature, instead relying on supervised models or multi-modal methods that align text and vision. We evaluate how the representations learned by modern self-supervised learning methods for vision perform under typical retrieval stacks that leverage vector databases and nearest neighbor search. Our evaluation reveals that the latent space geometry impacts approximate nearest neighbor (ANN) indexing. Specifically, highly anisotropic representations with high skewness produced by several modern SSL methods degrade the performance of partition-based and hashing-based search, even if their own linear probe or K-NN accuracy is not affected. In contrast, representations with higher isotropy and local purity better satisfy the distance-based assumptions of ANN indexes, leading to improved semantic retrieval performance.

自监督学习图像检索表征几何向量数据库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。