arXiv:2505.11060cs.CVcs.AI2025-05中稿 · IJCNN 2025被引 1

用视觉语言模型自动发现图像分类中的隐性偏见概念。

CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs

  • 基于图像文本潜在空间和线性探测器,无监督识别偏见概念。
  • 在不依赖失败样本或先验偏见知识下,成功发现未知偏见。
  • 适合研究模型公平性、追求可解释性的研究人员使用。

深度视觉模型常因数据集中存在的虚假相关性而习得偏见。现有基于概念的方法虽比热图等低层特征更易理解,但受限于缺乏标注的偏见概念数据,而人工标注成本高昂。本文提出CUBIC(Concept embeddings for Unsupervised Bias IdentifiCation),一种无需预设偏见候选或特定失败样本的无监督方法。该方法利用视觉语言模型(VLMs)的图像-文本潜在空间,通过线性分类器探测器分析超类标签的潜在表示如何受特定概念影响。通过测量这些变化与分类器决策边界法向量的夹角,识别显著影响模型预测的概念。实验表明,CUBIC能在未提供性能下降样本或先验偏见信息的情况下,有效发现新出现的偏见概念。

原文摘要 · Abstract (English)

Deep vision models often rely on biases learned from spurious correlations in datasets. To identify these biases, methods that interpret high-level, human-understandable concepts are more effective than those relying primarily on low-level features like heatmaps. A major challenge for these concept-based methods is the lack of image annotations indicating potentially bias-inducing concepts, since creating such annotations requires detailed labeling for each dataset and concept, which is highly labor-intensive. We present CUBIC (Concept embeddings for Unsupervised Bias IdentifiCation), a novel method that automatically discovers interpretable concepts that may bias classifier behavior. Unlike existing approaches, CUBIC does not rely on predefined bias candidates or examples of model failures tied to specific biases, as such information is not always available. Instead, it leverages image-text latent space and linear classifier probes to examine how the latent representation of a superclass label$\unicode{x2014}$shared by all instances in the dataset$\unicode{x2014}$is influenced by the presence of a given concept. By measuring these shifts against the normal vector to the classifier's decision boundary, CUBIC identifies concepts that significantly influence model predictions. Our experiments demonstrate that CUBIC effectively uncovers previously unknown biases using Vision-Language Models (VLMs) without requiring the samples in the dataset where the classifier underperforms or prior knowledge of potential biases.

偏见检测视觉语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。