提出GRAS基准与评分,量化视觉语言模型在性别种族等维度的偏见。
Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone
- 构建覆盖性别、种族、年龄、肤色的多维偏见评测基准GRAS。
- 五款顶尖模型中最优者偏见得分仅2/100,显示普遍严重偏见。
- 揭示提问方式影响评估结果,提醒需多角度设计测试问题。
随着视觉语言模型(VLMs)在现实应用中日益重要,理解其在人口统计学维度上的偏见至关重要。本文提出GRAS基准,用于检测VLMs在性别、种族、年龄和皮肤色调方面的偏见,覆盖范围为当前最广。我们进一步提出可解释的GRAS偏见评分,用于量化偏见程度。对五款先进VLMs进行测评,结果显示偏见水平令人担忧,最优模型的GRAS偏见评分仅为2/100。研究还揭示一个方法论洞见:在使用视觉问答(VQA)评估偏见时,必须考虑同一问题的不同表述形式。代码、数据及评估结果均已公开。
原文摘要 · Abstract (English)
As Vision Language Models (VLMs) become integral to real-world applications, understanding their demographic biases is critical. We introduce GRAS, a benchmark for uncovering demographic biases in VLMs across gender, race, age, and skin tone, offering the most diverse coverage to date. We further propose the GRAS Bias Score, an interpretable metric for quantifying bias. We benchmark five state-of-the-art VLMs and reveal concerning bias levels, with the least biased model attaining a GRAS Bias Score of only 2 out of 100. Our findings also reveal a methodological insight: evaluating bias in VLMs with visual question answering (VQA) requires considering multiple formulations of a question. Our code, data, and evaluation results are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。