测试视觉语言模型在新闻图中的偏见,发现性别和职业最易引发刻板印象。
Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
- 用1343张新闻图构建评测集,标注年龄、性别等社会属性
- 模型在开放任务中输出受视觉线索影响,性别/职业偏见尤其严重
- 高准确率模型仍可能有高偏见,适合关注公平性的研究者使用
大型视觉语言模型(VLMs)能联合理解图像与文本,但当图像包含年龄、性别、种族、着装或职业等视觉线索时,也容易吸收并再现有害社会刻板印象。为探究此类风险,我们构建了一个由1,343个图像-问题对组成的新闻图像基准数据集,涵盖多元媒体来源,并标注了真实答案及人口统计属性(年龄、性别、种族、职业、运动)。我们评估了多种前沿VLMs,并采用大语言模型(LLM)作为评判者,辅以人工验证。结果表明:(i)视觉上下文在开放式设置中系统性地改变模型输出;(ii)偏见程度随属性与模型而异,性别与职业存在显著风险;(iii)更高的忠实度并不意味着更低的偏见。我们已公开基准提示、评估标准与代码,支持可复现的多模态公平性评估。
原文摘要 · Abstract (English)
Large vision-language models (VLMs) can jointly interpret images and text, but they are also prone to absorbing and reproducing harmful social stereotypes when visual cues such as age, gender, race, clothing, or occupation are present. To investigate these risks, we introduce a news-image benchmark consisting of 1,343 image-question pairs drawn from diverse outlets, which we annotated with ground-truth answers and demographic attributes (age, gender, race, occupation, and sports). We evaluate a range of state-of-the-art VLMs and employ a large language model (LLM) as judge, with human verification. Our findings show that: (i) visual context systematically shifts model outputs in open-ended settings; (ii) bias prevalence varies across attributes and models, with particularly high risk for gender and occupation; and (iii) higher faithfulness does not necessarily correspond to lower bias. We release the benchmark prompts, evaluation rubric, and code to support reproducible and fairness-aware multimodal assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。