提出新基准与去偏方法,提升视觉语言模型公平性
debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias
- 构建更严谨的评估数据集,识别模型在性别、种族等维度的偏差
- 基于稀疏自编码器的去偏方法使模型公平性提升5-15点
- 发现图像类数据集比场景/代词类更适合作为偏见检测工具
随着视觉语言模型(VLMs)广泛应用,其公平性仍缺乏系统研究。本文分析了五种模型在六大数据集上的群体偏见,发现人脸数据集如UTKFace和CelebA是检测偏见的最佳工具,而场景类数据集(PATA、VLStereoSet)因文本提示可让模型猜答案,难以有效衡量偏见;代词类数据集(VisoGender)仅部分子集有参考价值。为此,我们提出更严格的评估数据集及基于稀疏自编码器的去偏方法。实验表明,新数据集能产生更具意义的错误信号,该方法使模型公平性提升5-15点,优于基线。本研究揭示了现有偏见评估基准的局限,并提供了更有效的测评工具与可解释的去偏方案。
原文摘要 · Abstract (English)
As Vision Language Models (VLMs) gain widespread use, their fairness remains under-explored. In this paper, we analyze demographic biases across five models and six datasets. We find that portrait datasets like UTKFace and CelebA are the best tools for bias detection, finding gaps in performance and fairness for both LLaVa and CLIP models. Scene-based datasets like PATA and VLStereoSet fail to be useful benchmarks for bias due to their text prompts allowing the model to guess the answer without a picture. As for pronoun-based datasets like VisoGender, we receive mixed signals as only some subsets of the data are useful in providing insights. To alleviate these two problems, we introduce a more rigorous evaluation dataset and a debiasing method based on Sparse Autoencoders to help reduce bias in models. We find that our data set generates more meaningful errors than the previous data sets. Furthermore, our debiasing method improves fairness, gaining 5-15 points in performance over the baseline. This study displays the problems with the current benchmarks for measuring demographic bias in Vision Language Models and introduces both a more effective dataset for measuring bias and a novel and interpretable debiasing method based on Sparse Autoencoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。