arXiv:2602.24014cs.CVcs.AI2026-02中稿 · CVPR被引 5

通过定位视觉语言模型中的社会属性神经元,实现可解释的公平性提升。

Interpretable Debiasing of Vision-Language Models for Social Fairness

  • 用稀疏自编码器在多模态编码器中识别出对特定群体敏感的神经元。
  • 仅关闭与偏见强相关的神经元,就能降低模型的社会偏见,且不损失语义理解能力。
  • 适合关注AI公平性、可解释性的研究人员和开发者使用。

视觉语言模型的快速发展引发了对其黑箱推理过程可能引发社会偏见的担忧。现有去偏方法主要依赖事后学习或测试时算法缓解表面偏见,却未深入探索模型内部机制。本文提出一种可解释、模型无关的去偏框架DeBiasLens,通过在多模态编码器上应用稀疏自编码器(SAEs)来定位社会属性神经元。在无社会属性标签的面部图像或文本数据集上训练SAEs,以发现对特定人口统计特征(包括少数群体)高度响应的神经元。通过选择性关闭每个群体中与偏见关联最强的神经元,有效缓解了视觉语言模型的社会偏见行为,同时保持其语义知识不受损害。本研究为未来审计工具奠定了基础,推动真实世界AI系统中的社会公平性建设。

原文摘要 · Abstract (English)

The rapid advancement of Vision-Language models (VLMs) has raised growing concerns that their black-box reasoning processes could lead to unintended forms of social bias. Current debiasing approaches focus on mitigating surface-level bias signals through post-hoc learning or test-time algorithms, while leaving the internal dynamics of the model largely unexplored. In this work, we introduce an interpretable, model-agnostic bias mitigation framework, DeBiasLens, that localizes social attribute neurons in VLMs through sparse autoencoders (SAEs) applied to multimodal encoders. Building upon the disentanglement ability of SAEs, we train them on facial image or caption datasets without corresponding social attribute labels to uncover neurons highly responsive to specific demographics, including those that are underrepresented. By selectively deactivating the social neurons most strongly tied to bias for each group, we effectively mitigate socially biased behaviors of VLMs without degrading their semantic knowledge. Our research lays the groundwork for future auditing tools, prioritizing social fairness in emerging real-world AI systems.

去偏可解释性视觉语言模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。