构建首个面向可持续发展目标的多模态偏见评测基准,揭示并缓解视觉语言模型的系统性偏差。
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

- 设计涵盖50万道选择题与5万回归任务的综合评测体系。
- 发现当前模型依赖特定目标先验而非多模态证据,导致预测偏差。
- 提出无需训练的CADE方法,提升准确率最多25%,降低误差12点。
评估可持续发展目标(SDGs)进展需对视觉线索、上下文知识和发展指标进行多步推理,不完整证据使用和不完善证据整合可能引入隐性预测偏差。现实中的SDG监测同时涉及定性判断与定量估计。然而,现有评测通常孤立评估这些方面,掩盖了模型用先验替代证据时产生的系统性偏差。为此,我们提出SDGBiasBench,一个大规模面向SDG的视觉-语言推理评测基准。该基准包含50万道专家参与的多选题和5万项回归任务,可全面评估视觉-语言模型(VLMs)在决策与估计层面的偏见。在该基准上的评估显示,当前VLMs存在内在的SDG偏见,其预测常由特定目标先验驱动,而非可靠多模态线索。为缓解此问题,我们提出CADE(对比自适应去偏集成),一种无需训练、即插即用的方法,利用模态特异性答案先验。CADE在多个VLM上显著提升性能,多选题准确率最高提升25%,回归任务平均绝对误差(MAE)降低最多12点。希望本工作能推动更公平、可靠的可持续发展人工智能系统发展。
原文摘要 · Abstract (English)
Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplete evidence use and imperfect evidence integration can introduce hidden prediction biases. Real-world SDG monitoring further spans both qualitative judgments and quantitative estimation. However, existing benchmarks typically evaluate these aspects in isolation, obscuring systematic biases that emerge when models substitute priors for evidence. To address this gap, we propose SDGBiasBench, a large-scale benchmark suite for SDG-oriented vision-language reasoning. Spanning 500k expert-involved multiple-choice questions and 50k regression tasks, the benchmark enables comprehensive assessment of both decision-level and estimation-level bias in Vision--Language Models (VLMs). Evaluations on SDGBiasBench reveal an intrinsic SDG bias in current VLMs, where predictions are frequently driven by SDG specific priors rather than reliable multi-modal cues. To mitigate such bias, we propose CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors. CADE yields significant gains on the proposed benchmark, improving multiple-choice accuracy by up to 25% and reducing regression MAE by up to 12 points across multiple VLMs. We hope our work can foster the development of more fair and reliable AI systems for sustainable development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。