arXiv:2410.19314cs.CYcs.CL2024-10ICLR被引 15

22个视觉语言模型存在性别偏见,可借微调有效缓解

Revealing and Reducing Gender Biases in Vision and Language Assistants (VLAs)

  • 分析22个开源视觉语言模型的性别偏见
  • 模型倾向给女性更多技能和正面性格,男性关联负面特质
  • 微调方法在去偏与性能间平衡最佳,适合部署前评估

预训练大语言模型已广泛与视觉输入结合用于多模态任务。指令微调的图像到文本视觉语言助手(VLAs)如LLaVA和InternVL的普及,亟需评估其性别偏见。我们研究了22个主流开源VLAs在人格特质、技能和职业方面的性别偏见。结果表明,这些模型复现了数据中可能存在的现实职业失衡等人类偏见。同时,模型倾向于赋予女性更多技能和积极人格特质,而对男性则有更一致的负面特质关联。为消除此类偏见,我们发现基于微调的去偏方法在去偏效果与下游任务性能保持之间取得最佳平衡。我们主张在部署前进行性别偏见评估,并推动去偏策略发展,以确保社会公平。代码开源于https://github.com/ExplainableML/vla-gender-bias。

原文摘要 · Abstract (English)

Pre-trained large language models (LLMs) have been reliably integrated with visual input for multimodal tasks. The widespread adoption of instruction-tuned image-to-text vision-language assistants (VLAs) like LLaVA and InternVL necessitates evaluating gender biases. We study gender bias in 22 popular open-source VLAs with respect to personality traits, skills, and occupations. Our results show that VLAs replicate human biases likely present in the data, such as real-world occupational imbalances. Similarly, they tend to attribute more skills and positive personality traits to women than to men, and we see a consistent tendency to associate negative personality traits with men. To eliminate the gender bias in these models, we find that fine-tuning-based debiasing methods achieve the best trade-off between debiasing and retaining performance on downstream tasks. We argue for pre-deploying gender bias assessment in VLAs and motivate further development of debiasing strategies to ensure equitable societal outcomes. Code is available at https://github.com/ExplainableML/vla-gender-bias.

视觉语言模型性别偏见去偏多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。