arXiv:2508.07432cs.CVcs.AI2025-08被引 2

揭示视觉语言模型的性别偏见来源,提出高效去偏方法

Freeze and Reveal: Exposing Modality Bias in Vision-Language Models

  • 用反事实数据增强和任务向量分离视觉与文本模态的偏见影响
  • 新方法DAUDoS仅用1/3数据降低3%偏见,识别准确率提升3%
  • 发现CLIP视觉模块更偏,PaliGemma2文本模块更偏,利于精准干预

视觉语言模型虽在多模态任务中表现优异,却常继承训练数据中的性别偏见,这些偏见可能来自视觉或文本模态。本文通过对抗性去偏方法,分别评估视觉与文本骨干网络对偏见的贡献。受仇恨言论分类中数据高效方法启发,提出新的刻板印象度量指标与去偏方法DAUDoS,以极低计算成本减少偏见。我们构建了带性别标注的数据集,并在VisoGender基准上评估所有方法。结果表明,CDA使性别差距缩小6%,DAUDoS缩小3%,且仅需1/3数据;两种方法均使图像性别识别准确率提升3%。实验显示,CLIP的视觉编码器偏见更强,而PaliGemma2的文本编码器偏见更显著。该研究为未来多模态系统提供针对性去偏策略。

原文摘要 · Abstract (English)

Vision Language Models achieve impressive multi-modal performance but often inherit gender biases from their training data. This bias might be coming from both the vision and text modalities. In this work, we dissect the contributions of vision and text backbones to these biases by applying targeted debiasing using Counterfactual Data Augmentation and Task Vector methods. Inspired by data-efficient approaches in hate-speech classification, we introduce a novel metric, Degree of Stereotypicality and a corresponding debiasing method, Data Augmentation Using Degree of Stereotypicality - DAUDoS, to reduce bias with minimal computational cost. We curate a gender annotated dataset and evaluate all methods on VisoGender benchmark to quantify improvements and identify dominant source of bias. Our results show that CDA reduces the gender gap by 6% and DAUDoS by 3% but using only one-third of the data. Both methods also improve the model's ability to correctly identify gender in images by 3%, with DAUDoS achieving this improvement using only almost one-third of training data. From our experiment's, we observed that CLIP's vision encoder is more biased whereas PaliGemma2's text encoder is more biased. By identifying whether bias stems more from vision or text encoders, our work enables more targeted and effective bias mitigation strategies in future multi-modal systems.

视觉语言模型偏见检测去偏方法多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。