评测大模型在真实图像中的视觉刻板印象,发现顶尖模型仍存偏见。
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
- 用14,144组真实多角色图像+问题对评估模型
- 19个主流模型在9类刻板印象上普遍表现偏差
- 揭示推理模型的思维链会放大偏见,适合公平性研究者使用
大型多模态模型(LMMs)中的刻板印象偏见会加剧社会偏见,威胁AI应用的公平性。随着LMMs影响力扩大,解决其在现实场景中对刻板印象、有害生成和模糊假设的固有偏见变得至关重要。然而现有数据集缺乏多样性,依赖合成图像且多为单角色图像,难以评估真实视觉情境下的偏见。为此,我们提出BBQ-Vision(BBQ-V),首个涵盖九类共50个子类别的综合性基准,采用真实多角色图像。该基准包含14,144个图像-问题对,通过精心设计的视觉基础场景,挑战模型对视觉刻板印象的准确推理能力。它提供基于真实视觉样本、图像变体及开放问答格式的稳健评估框架,可精确衡量模型在不同难度下的推理表现。通过对19个先进开源与闭源LMMs的严格测试,我们发现这些顶尖模型在多个社会刻板印象上仍存在显著偏差,并揭示推理型模型在思维链中会引入更多偏见。本工作推动了AI公平性发展,为更公平、负责任的LMMs奠定基础。数据集与代码已公开。
原文摘要 · Abstract (English)
Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influential, addressing and mitigating inherent biases related to stereotypes, harmful generations, and ambiguous assumptions in real-world scenarios has become essential. However, existing datasets evaluating stereotype biases in LMMs often lack diversity, rely on synthetic images, and often have single-actor images, leaving a gap in bias evaluation for real-world visual contexts. To address the gap in bias evaluation using real images, we introduce the BBQ-Vision (BBQ-V), the most comprehensive framework for assessing stereotype biases across nine diverse categories and 50 sub-categories with real and multi-actor images. BBQ-V benchmark contains 14,144 image-question pairs and rigorously evaluates LMMs through carefully curated, visually grounded scenarios, challenging them to reason accurately about visual stereotypes. It offers a robust evaluation framework featuring real-world visual samples, image variations, and open-ended question formats. BBQ-V enables a precise and nuanced assessment of a model's reasoning capabilities across varying levels of difficulty. Through rigorous testing of 19 state-of-the-art open-source (general-purpose and reasoning) and closed-source LMMs, we highlight that these top-performing models are often biased on several social stereotypes, and demonstrate that the thinking models induce more bias in the reasoning chains. This benchmark represents a significant step toward fostering fairness in AI systems and reducing harmful biases, laying the groundwork for more equitable and socially responsible LMMs. Our dataset and evaluation code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。