检测视觉语言模型对残障人士描述中的偏见与误判
Auditing Disability Representation in Vision-Language Models
- 用对照提示测试模型在残障语境下的描述准确性
- 9类残障中,引入残障背景后描述偏差率上升37%
- 适合关注AI公平性与伦理的开发者和研究者
视觉语言模型(VLMs)被广泛应用于社会敏感场景,但其对残障群体的表现仍缺乏深入研究。本文聚焦于以人物为中心的图像描述,发现模型常从基于视觉证据的客观描述转向包含未经证实推断的主观解读。为此,我们构建了一个基于中性提示(NP)与残障情境提示(DP)配对的基准,零样本评估了15个前沿开源及闭源VLM,在9类残障类别上进行测试。评估框架以解释保真度为核心目标,结合文本指标(情感、社会评价、响应长度)与由残障经历者验证的LLM裁判协议。结果表明,引入残障语境会显著降低解释保真度,引发推测性推断、叙事扩展、情感恶化及缺陷导向表述;这些影响在种族与性别维度上进一步加剧。最后,我们证明针对性提示与偏好微调可有效提升解释保真度,大幅减少解释偏移。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly deployed in socially sensitive applications, yet their behavior with respect to disability remains underexplored. We study disability aware descriptions for person centric images, where models often transition from evidence grounded factual description to interpretation shift including introduction of unsupported inferences beyond observable visual evidence. To systematically analyze this phenomenon, we introduce a benchmark based on paired Neutral Prompts (NP) and Disability-Contextualised Prompts (DP) and evaluate 15 state-of-the-art open- and closed-source VLMs under a zero-shot setting across 9 disability categories. Our evaluation framework treats interpretive fidelity as core objective and combines standard text-based metrics capturing affective degradation through shifts in sentiment, social regard and response length with an LLM-as-judge protocol, validated by annotators with lived experience of disability. We find that introducing disability context consistently degrades interpretive fidelity, inducing interpretation shifts characterised by speculative inference, narrative elaboration, affective degradation and deficit oriented framing. These effects are further amplified along race and gender dimension. Finally, we demonstrate targeted prompting and preference fine-tuning effectively improves interpretive fidelity and reduces substantially interpretation shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。