梳理文本生成图像模型的偏见与公平性,提出可落地的评估框架。
Operationalizing Fairness in Text-to-Image Models: A Survey of Bias, Fairness Audits and Mitigation Strategies

- 构建偏见类型与公平性概念的分类体系,厘清术语模糊问题。
- 揭示理想公平与可执行标准间的差距,强调可操作性重要性。
- 提出基于目标的测试框架,助力负责任生成式AI研发。
文本到图像(T2I)生成模型在各行业广泛应用,但常呈现社会刻板印象而受批评。尽管相关研究不断涌现,评估与缓解偏见的方法仍面临概念模糊问题,如‘偏见’与‘公平性’缺乏明确定义。本文系统综述了T2I公平性研究,构建偏见类型与公平性理念的分类体系。我们批判性分析‘目标公平性’(输出的理想规范)与‘阈值公平性’(具有可行动决策规则的标准)之间的差距。进一步调研从提示工程到扩散过程调控等缓解策略。最后提出一种新框架,推动公平性从描述性度量迈向严格的目标导向测试,为更可问责的生成式AI开发提供路径。
原文摘要 · Abstract (English)
Text-to-Image (T2I) generation models have been widely adopted across various industries, yet are criticized for frequently exhibiting societal stereotypes. While a growing body of research has emerged to evaluate and mitigate these biases, the field at present contends with conceptual ambiguity, for example terms like "bias" and "fairness" are not always clearly distinguished and often lack clear operational definitions. This paper provides a comprehensive systematic review of T2I fairness literature, organizing existing work into a taxonomy of bias types and fairness notions. We critically assess the gap between "target fairness" (normative ideals in T2I outputs) and "threshold fairness" (normative standards with actionable decision rules). Furthermore, we survey the landscape of mitigation strategies, ranging from prompt engineering to diffusion process manipulation. We conclude by proposing a new framework for operationalizing fairness that moves beyond descriptive metrics towards rigorous, target-based testing, offering an approach for more accountable generative AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。