arXiv:2410.13976cs.CVcs.CL2024-10被引 6

通过消除视觉语言模型中的敏感属性表示,实现无训练去偏。

Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations

  • 直接在生成阶段消除敏感属性表征,无需重新训练。
  • 仅需约1000个带偏见样本即可有效减少敏感属性提及。
  • 去偏后仍保持原有图像描述性能,适合实际应用部署。

大型视觉语言模型(如LLaVA)虽具备强大的多模态对话能力,但其响应易受训练数据中社会偏见影响,导致对不同人口群体的图像产生差异化输出。本文提出一种新型去偏框架,通过在文本生成过程中直接消除敏感属性表征,避免生成与受保护属性相关的文本或内部表示。该方法无需训练,仅需约1000个代表性偏见输出样本。实验表明,不仅能显著降低模型对敏感属性的提及倾向,还能利用合成数据指导消融过程,同时在真实数据集(如COCO)上维持良好的图像描述性能。此外,去偏后的生成结果在准确性上与基线有偏模型相当,证明去偏可不以牺牲性能为代价。

原文摘要 · Abstract (English)

Large Vision Language Models (LVLMs) such as LLaVA have demonstrated impressive capabilities as general-purpose chatbots that can engage in conversations about a provided input image. However, their responses are influenced by societal biases present in their training datasets, leading to undesirable differences in how the model responds when presented with images depicting people of different demographics. In this work, we propose a novel debiasing framework for LVLMs by directly ablating biased attributes during text generation to avoid generating text related to protected attributes, or even representing them internally. Our method requires no training and a relatively small amount of representative biased outputs (~1000 samples). Our experiments show that not only can we can minimize the propensity of LVLMs to generate text related to protected attributes, but we can even use synthetic data to inform the ablation while retaining captioning performance on real data such as COCO. Furthermore, we find the resulting generations from a debiased LVLM exhibit similar accuracy as a baseline biased model, showing that debiasing effects can be achieved without sacrificing model performance.

去偏视觉语言模型敏感属性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。