arXiv:2508.03079cs.CV2025-08被引 2

通过反事实VQA检测闭源模型的非人口统计偏差,发现环境与行为因素影响更大。

Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA

  • 设计反事实VQA基准,仅改变一个无关视觉属性测试决策稳定性。
  • 非人口统计属性导致的偏差强度超过人口统计属性。
  • 少量人类规范示例可提升模型响应一致性,适合模型审计与优化。

大型视觉语言模型(LVLMs)的进展加剧了公平性担忧,但现有评估仍局限于人口统计属性,且常将公平性与拒绝行为混淆。本文提出一种反事实VQA基准,通过受控上下文变化探测闭源LVLM的决策边界。每组图像仅在一项经验证无关的视觉属性上差异,实现无真实答案依赖且能识别拒绝行为的推理稳定性分析。全面实验表明,环境上下文或社会行为等非人口统计属性对LVLM决策的影响强于人口统计属性。此外,基于指令的去偏方法效果有限,甚至可能加剧不对称性;而引入少量由本基准中人类规范验证的示例,可促使模型产生更一致、更平衡的响应,凸显该基准不仅可用作评估工具,还可用于理解与改进模型行为。这些结果为审计黑箱LVLM中的情境偏差提供了实践基础,推动多模态推理的透明与公平。

原文摘要 · Abstract (English)

Recent advances in large vision-language models (LVLMs) have amplified concerns about fairness, yet existing evaluations remain confined to demographic attributes and often conflate fairness with refusal behavior. This paper broadens the scope of fairness by introducing a counterfactual VQA benchmark that probes the decision boundaries of closed-source LVLMs under controlled context shifts. Each image pair differs in a single visual attribute that has been validated as irrelevant to the question, enabling ground-truth-free and refusal-aware analysis of reasoning stability. Comprehensive experiments reveal that non-demographic attributes, such as environmental context or social behavior, distort LVLM decision-making more strongly than demographic ones. Moreover, instruction-based debiasing shows limited effectiveness and can even amplify these asymmetries, whereas exposure to a small number of human norm validated examples from our benchmark encourages more consistent and balanced responses, highlighting its potential not only as an evaluative framework but also as a means for understanding and improving model behavior. Together, these results provide a practial basis for auditing contextual biases even in black-box LVLMs and contribute to more transparent and equitable multimodal reasoning.

视觉语言模型公平性反事实推理去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。