提出几何化去偏方法,将偏见视为子空间而非坐标,更彻底地消除视觉语言模型中的性别种族偏见。
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
- 将偏见视为线性子空间,整体投影移除而非替换单个坐标
- 在4个公平性指标上平均提升18.5%,任务性能损失极小
- 适用于需高公平性的图像分类、图文检索与生成任务
视觉语言模型在多模态推理中不可或缺,但其表征常编码并放大人口统计学偏见,导致下游任务出现偏差关联与预测失准,损害公平性并扭曲视觉与语言的对齐。现有后处理去偏方法通过替换最相关属性的嵌入坐标实现,但系统分析揭示其存在特征纠缠、跨数据集泛化差、偏见清除不彻底三大缺陷。我们发现偏见并非集中于少数坐标,而是分布于若干线性子空间。为此,提出子空间投影去偏(SPD),一种基于几何原理的框架:识别并移除可线性解码的偏见子空间,并重新插入中性均值成分以保持语义保真。在零样本分类、文本到图像检索与图像生成任务上的大量实验验证其有效性:相比最佳基线,平均提升18.5%的公平性指标,同时任务性能损失最小。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have become indispensable for multimodal reasoning, yet their representations often encode and amplify demographic biases, resulting in biased associations and misaligned predictions in downstream tasks. Such behavior undermines fairness and distorts the intended alignment between vision and language. Recent post-hoc approaches attempt to mitigate bias by replacing the most attribute-correlated embedding coordinates with neutral values. However, our systematic analysis reveals three critical limitations of this coordinate-wise approach: feature entanglement, poor cross-dataset generalization, and incomplete bias removal. We find that bias is not localized to a few coordinates but is instead distributed across a few linear subspaces. To address these limitations, we propose $\textbf{S}$ubspace $\textbf{P}$rojection $\textbf{D}$ebiasing ($\textbf{SPD}$), a geometrically principled framework that identifies and removes the entire subspace of linearly decodable bias while reinserting a neutral mean component to preserve semantic fidelity. Extensive experiments across zero-shot classification, text-to-image retrieval, and image generation validate the effectiveness of SPD: our method achieves more robust debiasing with an average improvement of $18.5\%$ across four fairness metrics, while maintaining minimal loss in task performance compared to the best debiasing baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。