首个综合评估文生图模型三类偏见的框架,揭示主流模型仍存显著性别与文化偏差。
T2I-BiasBench: A Multi-Metric Framework for Auditing Demographic and Cultural Bias in Text-to-Image Models
- 构建13项互补指标,统一评估性别、元素缺失与文化扁平化三类偏见。
- 发现稳定扩散等模型在美相关提示中放大偏见,且所有模型文化多样性不足。
- 适用于模型开发者、评测人员及政策制定者进行系统性偏见检测与改进。
文生图生成模型虽具备出色视觉质量,但继承并放大训练数据中的种族不平等与文化偏见。本文提出T2I-BiasBench,首个同时涵盖三大维度(性别偏见、元素缺失、文化坍缩)的统一评估框架,包含十三项互补指标:整合六项已有指标,并新增四项(综合偏见得分、有根基缺失率、隐含元素缺失率、文化准确率)及三项适配指标(幻觉得分、Vendi得分、CLIP代理得分)。在五类结构化提示下生成1,574张图像,评估了Stable Diffusion v1.5、BK-SDM Base、Koala Lightning三款开源模型,并以经强化学习人类反馈对齐的Gemini 2.5 Flash为基准。关键发现:(1) Stable Diffusion v1.5与BK-SDM在美相关提示中呈现偏见放大(>1.0);(2) 手术防护装备等上下文约束可显著降低职业角色性别偏见(医生类别的偏见综合得分CBS = 0.06);(3) 所有模型(包括对齐后的Gemini)均陷入狭窄文化表达(文化代表性分数CAS: 0.54–1.00),表明对齐无法解决文化覆盖缺口。T2I-BiasBench已公开发布,支持生成模型的标准化、细粒度偏见评估。
原文摘要 · Abstract (English)
Text-to-image (T2I) generative models achieve impressive visual fidelity but inherit and amplify demographic imbalances and cultural biases embedded in training data. We introduce T2I-BiasBench, a unified evaluation framework of thirteen complementary metrics that jointly captures demographic bias, element omission, and cultural collapse in diffusion models - the first framework to address all three dimensions simultaneously. We evaluate three open-source models - Stable Diffusion v1.5, BK-SDM Base, and Koala Lightning - against Gemini 2.5 Flash (RLHF-aligned) as a reference baseline. The benchmark comprises 1,574 generated images across five structured prompt categories. T2I-BiasBench integrates six established metrics with seven additional measures: four newly proposed (Composite Bias Score, Grounded Missing Rate, Implicit Element Missing Rate, Cultural Accuracy Ratio) and three adapted (Hallucination Score, Vendi Score, CLIP Proxy Score). Three key findings emerge: (1) Stable Diffusion v1.5 and BK-SDM exhibit bias amplification (>1.0) in beauty-related prompts; (2) contextual constraints such as surgical PPE substantially attenuate professional-role gender bias (Doctor CBS = 0.06 for SD v1.5); and (3) all models, including RLHF-aligned Gemini, collapse to a narrow set of cultural representations (CAS: 0.54-1.00), confirming that alignment techniques do not resolve cultural coverage gaps. T2I-BiasBench is publicly released to support standardized, fine-grained bias evaluation of generative models. The project page is available at: https://gyanendrachaubey.github.io/T2I-BiasBench/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。