arXiv:2512.00807cs.AI2025-12被引 1

提出训练免扰的性别公平框架,区分中性与明确场景下的性别信息处理。

BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

  • 通过反事实嵌入识别性别变异子空间,实现选择性去偏。
  • 中性场景下性别偏见降低37.6%,显式场景中性别特征保留率达92.1%。
  • 适用于连续偏见变量,可推广至亮度等多类偏见控制。

视觉语言模型(VLMs)继承了训练数据中的显著社会偏见,尤其在性别表征方面。现有公平干预通常采取无差别的统一处理视角,未能区分需保持中立与应保留群体特性的场景。基于文本模型中差异感知公平的进展,我们将其拓展至多模态领域,形式化图像字幕与文生图任务中的差异感知性别公平问题。倡导选择性去偏:在中性场景中缓解不当偏见,而在显式场景中保留合法的群体差异。为此,提出完全无需训练的BioPro(偏差正交投影)框架。BioPro通过反事实嵌入识别低维性别变异子空间,并应用投影实现选择性中和性别相关信息。实验表明,BioPro在中性情况下有效降低性别偏见,同时在显式场景中维持92.1%的性别忠实度。此外,该方法还能泛化至连续偏见变量(如场景亮度),展现更广泛应用潜力。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interventions often adopt a difference-unaware perspective that enforces uniform treatment across demographic groups. These approaches, however, fail to distinguish between contexts where neutrality is required and those where group-specific attributes are legitimate and must be preserved. Building upon recent advances in difference-aware fairness for text-only models, we extend this concept to the multimodal domain and formalize the problem of difference-aware gender fairness for image captioning and text-to-image generation. We advocate for selective debiasing, which aims to mitigate unwanted bias in neutral contexts while preserving valid distinctions in explicit ones. To achieve this, we propose BioPro (Bias Orthogonal Projection), an entirely training-free framework. BioPro identifies a low-dimensional gender-variation subspace through counterfactual embeddings and applies projection to selectively neutralize gender-related information. Experiments show that BioPro effectively reduces gender bias in neutral cases while maintaining gender faithfulness in explicit ones, thus providing a promising direction toward achieving selective fairness in VLMs. Beyond gender bias, we further demonstrate that BioPro can effectively generalize to continuous bias variables, such as scene brightness, highlighting its broader applicability.

视觉语言模型性别公平选择性去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。