arXiv:2506.09106cs.CVcs.LG2025-06

分析无条件图像生成模型中的偏见机制,发现评估结果受分类器影响大。

Bias Analysis in Unconditional Image Generative Models

  • 用概率差定义属性偏见,对比训练与生成分布。
  • 实验显示属性偏移量小,但结果受分类器敏感性影响显著。
  • 适合关注生成模型伦理、评估框架设计的研究者阅读。

生成式AI的广泛应用引发了对表征伤害和潜在歧视结果的担忧。尽管相关研究增多,但无条件生成中偏见的形成机制仍不清晰。本文将属性偏见定义为观测分布中属性出现概率与理想参考分布期望比例之间的差异。通过训练一系列无条件图像生成模型,并采用常见偏见评估框架,研究训练分布与生成分布间的偏见变化。实验表明,检测到的属性偏移较小。进一步发现,偏见评估结果对用于标注生成图像的属性分类器高度敏感,尤其当分类器决策边界位于高密度区域时。实证分析表明,这种敏感性在连续型属性(而非二元属性)上更为常见,凸显了标签实践需更代表性的必要性,需加强对评估框架的审视,并正视属性的社会复杂性。

原文摘要 · Abstract (English)

The widespread adoption of generative AI models has raised growing concerns about representational harm and potential discriminatory outcomes. Yet, despite growing literature on this topic, the mechanisms by which bias emerges - especially in unconditional generation - remain disentangled. We define the bias of an attribute as the difference between the probability of its presence in the observed distribution and its expected proportion in an ideal reference distribution. In our analysis, we train a set of unconditional image generative models and adopt a commonly used bias evaluation framework to study bias shift between training and generated distributions. Our experiments reveal that the detected attribute shifts are small. We find that the attribute shifts are sensitive to the attribute classifier used to label generated images in the evaluation framework, particularly when its decision boundaries fall in high-density regions. Our empirical analysis indicates that this classifier sensitivity is often observed in attributes values that lie on a spectrum, as opposed to exhibiting a binary nature. This highlights the need for more representative labeling practices, understanding the shortcomings through greater scrutiny of evaluation frameworks, and recognizing the socially complex nature of attributes when evaluating bias.

图像生成偏见分析评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。