评测视觉语言模型在招聘、法律、医疗中的公平性与可靠性,发现模型常错误推断身份特征。
FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making
- 构建包含超10万张人脸问答对的基准,从公平性、有效性等多角度评估模型表现
- 发现模型主要问题在于过度推断而非不平等对待,最差模型99%无法回答时仍强行推断
- 强调在高风险场景中拒绝无依据推断比统计公平更重要,适合关注AI伦理的研究者
视觉语言模型(VLMs)正越来越多地用于基于视觉输入的决策。我们提出FAIRLENS,一个针对招聘、法律和医疗三个高风险领域中VLM响应的公平性与有效性评估基准。FAIRLENS将涵盖性别、种族和年龄群体的真实人脸图像与封闭式及开放式问题配对,每种模型生成超过10万组图像-问题对。评估从四个维度展开:负面结果率的群体均等性、回答合理性、与无支持角色/状态的群体关联性,以及自由文本生成中的偏见。合理性是核心有效性标准:当图像不足以支持回答时,模型应选择不回答。评估8个VLM后发现,主要问题是不当推断而非不平等对待。模型经常从面部推断资质、威胁、疾病或职业角色,而不选择放弃;最弱模型在无法回答的问题上仍进行推断的比例高达99%。此类缺陷在法律与医疗领域尤为严重,此时识别证据不足至关重要。仅看统计差异会掩盖问题:尽管绝对差距小,但若基线不良率低,相同差距意味着某一群体被标注为负面的情况可能是另一群体的数倍;小差距也可能反映所有群体都存在不安全行为。自由文本偏见与多选题准确率关联较弱,结构化答案正确并不等于生成内容安全。FAIRLENS表明,高风险场景下的公平模型需在群体间保持一致处理,并拒绝根据外貌推断关键属性;其问题集可迁移至任何带人口统计标注的人脸数据集。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used to make decisions from visual inputs. We introduce FAIRLENS, a benchmark and evaluation framework for measuring both the fairness and the validity of VLM responses in three high-stakes domains: hiring, legal, and healthcare. FAIRLENS pairs real face images spanning gender, race, and age groups with closed- and open-ended questions, giving more than 100K image-question pairs per model, and evaluates responses from four complementary views: demographic parity over adverse outcome rates, soundness, demographic association over unsupported roles and statuses, and bias in free-text generation. Soundness is the central validity criterion: a response is sound when it follows the evidence stated in the question and abstains when the image cannot support an answer. Evaluating eight VLMs, we find that the primary failure is unwarranted inference rather than unequal treatment. Models routinely infer qualifications, threat, illness, or professional role from a face instead of abstaining, and the weakest model does so on 99% of the questions its input cannot answer. These failures are most severe in legal and healthcare, where recognizing insufficient evidence matters most, and disparity metrics alone would miss them: parity gaps are small in absolute terms, yet when baseline adverse rates are low the same gap means one demographic group receives adverse labels several times as often as another, and a small gap can equally reflect a model that treats every group unsafely. Bias in free-text responses is only loosely coupled to multiple-choice accuracy, so correct structured answers do not imply safe generation. FAIRLENS shows that fair high-stakes VLM behavior requires similar treatment across groups and refusal to infer high-stakes attributes from appearance, and its question suite transfers to any face corpus with demographic annotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。