发现视觉语言模型的公平性悖论并提出可解释的纠偏方法
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
- 通过分析隐藏状态动态揭示模型内部偏见机制
- 中间层公平知识最强,末层出现显著衰减
- 基于残差流调整实现推理校准,不损失推理能力
尽管视觉语言模型(VLMs)取得了显著进展,其社会偏见背后的知识机制仍如黑箱般未知,导致特定群体受到不公平对待。本文系统研究了先进VLMs在生成响应中对性别与种族的偏见,不仅关注表面输出,还深入分析内部概率分布与隐藏状态动态。实证发现:1)公平性悖论——模型生成看似公平的文本标签,但对特定社会群体的置信度严重失准;2)层间波动——公平知识并非均匀分布,集中在中间层,最终层出现显著知识退化;3)残差差异——单个隐藏层内不同残差路径承载冲突的社会认知,部分强化公平,部分放大偏见。基于此,提出后处理框架RES-FAIR,通过定位并投影偏离偏见残差方向的隐藏状态,增强公平成分。在PAIRS与SocialCounterfactuals数据集上的评估表明,该方法显著提升响应公平性与置信度校准,且不损害通用推理能力。本工作为理解多模态模型如何存储与处理敏感社会信息提供了新视角。
原文摘要 · Abstract (English)
Although Vision-Language Models (VLMs) have achieved remarkable success, the knowledge mechanisms underlying their social biases remain a black box, where fairness- and ethics-related problems harm certain groups of people in society. It is unknown to what extent VLMs yield gender and race bias in generative responses. In this paper, we conduct a systematic discovery of gender and race bias in state-of-the-art VLMs, focusing not only on surface-level responses but also on the internal probability distributions and hidden state dynamics. Our empirical analysis reveals three critical findings: 1) The Fairness Paradox: Models often generate fair text labels while maintaining highly skewed confidence scores (mis-calibration) toward specific social groups. 2) Layer-wise Fluctuation: Fairness knowledge is not uniformly distributed; it peaks in intermediate layers and undergoes substantial knowledge erosion in the final layers. 3) Residual Discrepancy: Within a single hidden layer, different residual streams carry conflicting social knowledge - some reinforcing fairness while others amplifying bias. Leveraging these insights, we propose RES-FAIR (RESidual Flow Adjustment for Inference Recalibration), a post-hoc framework that mitigates bias by localizing and projecting hidden states away from biased residual directions while amplifying fair components. Evaluations on PAIRS and SocialCounterfactuals datasets demonstrate that our discovery-based approach significantly improves response fairness and confidence calibration without compromising general reasoning abilities. Our work provides a new lens for understanding how multi-modal models store and process sensitive social information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。