用注意力图检测模型在人脸数据中的隐性偏见
Attention IoU: Examining Biases in CelebA using Attention Maps
- 通过注意力交并比衡量模型内部表征的偏见
- 在CelebA上发现标签外的特征关联,揭示隐藏偏见
- 适合研究模型公平性与可解释性的研究人员
计算机视觉模型在多种数据集和任务中表现出并放大偏见。现有量化分类模型偏见的方法主要关注数据分布和子群体性能差异,忽视了模型内部机制。本文提出注意力交并比(Attention-IoU)度量及其相关得分,利用注意力图揭示模型内部表征中的偏见,并识别可能导致偏见的图像特征。首先在合成的Waterbirds数据集上验证Attention-IoU能准确测量模型偏见;随后分析CelebA数据集,发现Attention-IoU能捕捉超出准确率差异的关联。通过以性别为受保护属性,研究各属性的偏见表现方式。最后,通过子采样训练集改变属性相关性,证明Attention-IoU能揭示数据标签中未体现的潜在混杂变量。
原文摘要 · Abstract (English)
Computer vision models have been shown to exhibit and amplify biases across a wide array of datasets and tasks. Existing methods for quantifying bias in classification models primarily focus on dataset distribution and model performance on subgroups, overlooking the internal workings of a model. We introduce the Attention-IoU (Attention Intersection over Union) metric and related scores, which use attention maps to reveal biases within a model's internal representations and identify image features potentially causing the biases. First, we validate Attention-IoU on the synthetic Waterbirds dataset, showing that the metric accurately measures model bias. We then analyze the CelebA dataset, finding that Attention-IoU uncovers correlations beyond accuracy disparities. Through an investigation of individual attributes through the protected attribute of Male, we examine the distinct ways biases are represented in CelebA. Lastly, by subsampling the training set to change attribute correlations, we demonstrate that Attention-IoU reveals potential confounding variables not present in dataset labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。