首个系统评估文生图模型残疾偏见的基准,揭示模型总把残障人士与轮椅绑定。
Beyond wheelchairs and blindfolds: Investigating disability stereotypes in T2I models with INCLUDE-BENCH

- 设计11.9万组提示词,覆盖静态动态场景和多维度偏见
- 所有模型生成残障相关图像时,90%以上都用轮椅,多样性极低
- 提出真实世界偏见度量指标,发现模型复刻了社会刻板印象
文生图(T2I)模型存在社会偏见。以往研究主要聚焦性别、肤色和文化在有限职业关联中的表现,新兴评测基准也逐步纳入这些维度,但残疾议题仍被系统性忽视。现有评估方法常脱离社会学定义,难以准确衡量对残障人士(PWD)的表征伤害。为此,我们提出INCLUDE-BENCH,首个大规模评估T2I模型中残疾相关偏见的基准。该基准包含119,000张基于多维偏见设计提示词生成的图像,涵盖静态与动态语境。我们评估了15个开源与2个闭源模型。关键发现包括:(1)行动不便及默认残障提示在所有模型中均主导生成轮椅图像;(2)残障条件生成的图像一致性显著降低,多样性不足;(3)刻板形象与文本提示的对齐程度更强;(4)引入刻板印象内容模型(SCM)得分,证实T2I模型反映了现实世界中的刻板关联。
原文摘要 · Abstract (English)
Text-to-image (T2I) models have been shown to exhibit social biases. Prior work has mainly focused on gender, skin tone, and cultural representation within restricted occupational associations, and emerging benchmarks increasingly incorporate these dimensions. However, disability remains systematically underexplored. Current evaluation practices often fail to align with sociologically grounded definitions of stereotyping, limiting principled assessment of representational harms toward people with disabilities (PWD). To address this, we introduce INCLUDE-BENCH, the first large-scale benchmark for evaluating disability-related bias in T2I models. INCLUDE-BENCH comprises 119K generated images based on prompt design across multiple bias dimensions and both static and dynamic contexts. We evaluate 15 open-source and 2 closed-source models. Our key findings reveal that: (1) mobility-impaired and default disability prompts predominantly yield wheelchair depictions across all models; (2) disability-conditioned generations consistently exhibit less diversity; (3) stereotypical portrayals demonstrate stronger disability-text alignment; and (4) we introduce the Stereotype Content Model (SCM) Score, demonstrating that T2I models reflect real-world stereotypical associations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。