提出多维度公平性评估与去偏框架,提升文生图模型的公平性。
HoloFair: Unified T2I Fairness Evaluation and Fair-GRPO Debiasing

- 构建多属性分组偏差指数MGBI,量化多重社会属性偏差
- 在SD3.5-Medium上实现公平性显著提升且图像质量保持优异
- 基于强化学习的Fair-GRPO可有效缓解偏见,适合模型开发者使用
文生图(T2I)模型在视觉真实性和语义一致性方面取得显著进展,但仍常加剧和放大社会偏见。现有评估方法多局限于单一维度偏见,难以揭示深层社会语义层面的模型偏见。本文提出HoloFair,一个面向多维度人口统计偏见分析的综合性基准框架。基于大规模公平性数据集与SpaFreq(空间-频率)属性分类器,该框架设计了多属性分组偏差指数(MGBI),用于评估内在多样性与条件偏见。在评估之外,进一步提出基于强化学习的Fair-GRPO去偏方法,通过设计多目标奖励函数调整生成模型分布。例如,在SD3.5-Medium模型上的实验表明,Fair-GRPO在显著提升多维度公平性的同时保持高质量图像生成。我们还分析了潜在的奖励欺骗现象,并提出相应缓解策略。代码与数据集已公开于https://github.com/1059684669/HoloFair。
原文摘要 · Abstract (English)
Text-to-Image (T2I) models have made significant strides in visual realism and semantic consistency, yet they often perpetuate and amplify societal biases. Existing evaluation methods typically address only single-dimensional biases, lacking perspectives to uncover model biases at social-related deeper semantic levels. We introduce HoloFair, a comprehensive benchmark framework for multidimensional demographic bias analysis. Built upon our large-scale fairness-oriented dataset and the SpaFreq (Spatial-Frequency) attribute classifier, this framework proposes the Multi-attribute, Group-wise Bias Index (MGBI) metric, designed to assess both intrinsic diversity and conditional biases. Beyond evaluation, we further introduce Fair-GRPO, a reinforcement-learning-based debiasing method that alters the distribution of generative models through a designed multi-objective reward function. E.g., experiments on the SD3.5-Medium model demonstrate that Fair-GRPO significantly improves multidimensional fairness while maintaining high image quality. We also analyze potential reward hacking phenomena and provide corresponding mitigation strategies. Code and dataset are available at https://github.com/1059684669/HoloFair
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。