用视觉概念分解法自动识别数据集偏见,无需人工标注。
ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- 通过稀疏自编码器从视觉模型中提取可解释的视觉概念。
- 能发现并量化背景、物体共现等已知与未知偏见。
- 适合数据审计和模型鲁棒性评估,尤其适用于图像分类任务。
数据集偏见在机器学习中普遍存在,即某些概念在数据中过度集中,但缺乏精细属性标注时难以系统识别。本文提出ConceptScope,一种可扩展且自动化的视觉数据集分析框架,利用基于视觉基础模型表示训练的稀疏自编码器,发现并量化人类可理解的视觉概念。ConceptScope将概念分为目标、上下文和偏见三类,依据其语义相关性和与类别标签的统计相关性,实现类别级数据集表征、偏见识别与鲁棒性评估。通过与标注数据集对比,验证了其能捕捉物体、纹理、背景、面部特征、情绪和动作等多种概念。此外,概念激活产生的空间归因与语义有意义区域一致。ConceptScope可靠检测出已知偏见(如Waterbirds中的背景偏见),并发现此前未标注的共现模式(如ImageNet中的共现物体),为数据集审计和模型诊断提供了实用工具。
原文摘要 · Abstract (English)
Dataset bias, where data points are skewed to certain concepts, is ubiquitous in machine learning datasets. Yet, systematically identifying these biases is challenging without costly, fine-grained attribute annotations. We present ConceptScope, a scalable and automated framework for analyzing visual datasets by discovering and quantifying human-interpretable concepts using Sparse Autoencoders trained on representations from vision foundation models. ConceptScope categorizes concepts into target, context, and bias types based on their semantic relevance and statistical correlation to class labels, enabling class-level dataset characterization, bias identification, and robustness evaluation through concept-based subgrouping. We validate that ConceptScope captures a wide range of visual concepts, including objects, textures, backgrounds, facial attributes, emotions, and actions, through comparisons with annotated datasets. Furthermore, we show that concept activations produce spatial attributions that align with semantically meaningful image regions. ConceptScope reliably detects known biases (e.g., background bias in Waterbirds) and uncovers previously unannotated ones (e.g, co-occurring objects in ImageNet), offering a practical tool for dataset auditing and model diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。