用自动提取的概念向量解释图像相似性的原因
Explaining Image Similarity with Automatically Extracted Concept Activation Vectors

- 通过稀疏自编码器自动发现概念方向并扰动嵌入空间
- 概念重要性能线性恢复真实相似度分数
- 可定位相似原因,适合理解模型决策逻辑
图像相似性支撑众多计算机视觉应用,但为何两张图像获得高或低相似度评分却不明确。现有可解释方法多依赖梯度归因图提供局部解释,难以揭示嵌入空间中驱动相似性的全局因素(如纹理、形状、颜色)。我们提出一种模型与度量无关的框架,利用稀疏自编码器(SAEs)自动提取的概念激活向量(CAVs)解释图像相似性。给定图像对,沿发现的概念方向扰动其嵌入,并测量所选相似度函数的变化,从而获得概念重要性;对图像对生成概念归因图实现定位。该方法扩展至群体层面,解释一组图像间的相似性驱动因素,并引入示例检索(Exemplar Retrieval),以恢复具有相似解释原因的样本。实验表明,我们的隐空间扰动比像素空间基线更符合数据分布,且概念重要性能线性恢复真实相似度分数。定性结果进一步验证了方法在理解模型个体与群体相似性判断中的有效性。
原文摘要 · Abstract (English)
Image similarity underlies many computer vision applications, yet it is often unclear why two images receive a high or low similarity score. Existing explainability methods often rely on gradient-based attribution maps to provide local justifications for similarity. These approaches struggle to provide global insights into what specifically drives similarity in regions of an embedding space, such as texture, shape, or color. We introduce a model- and metric-agnostic framework that explains image similarity using Concept Activation Vectors (CAVs) extracted automatically via Sparse Autoencoders (SAEs). Given a pair of images, we perturb their embeddings along discovered concept directions and measure the resulting change in a chosen similarity function, yielding concept importances. For image pairs, we provide localization with concept attribution maps. We extend this procedure to group-level settings, explaining what drives similarity across a cluster of images rather than a single pair, and further, we introduce Exemplar Retrieval, aiming to recover samples with similar reasons contributing to similarity. Our experiments show that our latent perturbations are more faithful to the underlying data distribution than pixel-space baselines, and that concept importances linearly recover the true similarity score. Qualitative results further confirm the usefulness of our methods in understanding a model's individual and group similarity judgments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。