arXiv:2604.13305cs.CV2026-04中稿 · The IEEE/CVF Confe…被引 1

发现图像生成中的奖励模型会加剧性别种族偏见。

Bias at the End of the Score

  • 用大规模审计验证奖励模型在生成中存在人口统计偏见。
  • 奖励引导优化使女性主体被过度性化,固化性别种族刻板印象。
  • 适合关注AI伦理、图像生成公平性的研究者阅读。

奖励模型(RMs)是为特定目标(如人类偏好或图文对齐)设计的非中立价值函数,已成为文本到图像(T2I)生成系统的关键组件,用于数据集过滤、评估、参数优化监督及生成后安全质量筛选。尽管已有研究关注其集成问题(如奖励劫持或模式崩溃),但其作为评分函数的鲁棒性与公平性仍不明确。本文开展大规模审计,评估在T2I训练与生成过程中奖励模型对人口统计偏见的敏感性。定量与定性证据表明,原本作为质量度量的奖励模型实际上编码了人口统计偏见,导致奖励引导优化对女性图像主体过度性化,强化性别与种族刻板印象,并降低群体多样性。这些发现揭示了当前奖励模型的缺陷,挑战其作为质量指标的可靠性,并强调需改进数据收集与训练流程以实现更稳健的评分。

原文摘要 · Abstract (English)

Reward models (RMs) are inherently non-neutral value functions designed and trained to encode specific objectives, such as human preferences or text-image alignment. RMs have become crucial components of text-to-image (T2I) generation systems where they are used at various stages for dataset filtering, as evaluation metrics, as a supervisory signal during optimization of parameters, and for post-generation safety and quality filtering of T2I outputs. While specific problems with the integration of RMs into the T2I pipeline have been studied (e.g. reward hacking or mode collapse), their robustness and fairness as scoring functions remains largely unknown. We conduct a large scale audit of RM robustness with respect to demographic biases during T2I model training and generation. We provide quantitative and qualitative evidence that while originally developed as quality measures, RMs encode demographic biases, which cause reward-guided optimization to disproportionately sexualize female image subjects reinforce gender/racial stereotypes, and collapse demographic diversity. These findings highlight shortcomings in current reward models, challenge their reliability as quality metrics, and underscore the need for improved data collection and training procedures to enable more robust scoring.

图像生成公平性奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。