13个主流图文模型生成图像时普遍放大性别刻板印象,且表现差异巨大。
Automated Evaluation of Gender Bias Across 13 Large Multimodal Models
- 用75个中性提示生成人物图像,通过大模型评分系统量化性别偏见。
- 男性在男性化职业中占比93%,女性在女性化职业中仅22.5%,非刻板职业也偏男性(68.3%)。
- 不同模型偏差程度悬殊,顶尖模型接近性别平等,说明偏见可被设计避免。
大型多模态模型(LMMs)虽革新了文生图技术,却可能延续训练数据中的社会偏见。现有研究虽发现性别偏见,但方法学限制阻碍了大规模、跨模型的可比分析。为此,我们提出Aymara Image Fairness Evaluation基准,评估AI生成图像的社会偏见。测试13个商业LMMs,使用75个程序生成的性别中性提示,生成从事典型男性、典型女性及非刻板职业的人物图像,并通过经验证的LLM-as-a-judge系统对965张图像进行性别表征评分。结果表明(所有p < .001):1)LMMs不仅复现甚至放大职业性别刻板印象——男性化职业中男性占比达93.0%,女性化职业中仅22.5%;2)模型存在显著默认男性偏见,在非刻板职业中男性出现率达68.3%;3)偏见程度模型间差异极大,整体男性占比在46.7%至73.3%之间。值得注意的是,表现最佳模型削弱了性别刻板印象,接近性别均衡,获得最高公平分。这一差异表明高偏见并非必然,而是设计选择的结果。本研究提供迄今最全面的跨模型性别偏见基准,强调建立标准化、自动化评估工具对推动AI开发责任与公平性的必要性。
原文摘要 · Abstract (English)
Large multimodal models (LMMs) have revolutionized text-to-image generation, but they risk perpetuating the harmful social biases in their training data. Prior work has identified gender bias in these models, but methodological limitations prevented large-scale, comparable, cross-model analysis. To address this gap, we introduce the Aymara Image Fairness Evaluation, a benchmark for assessing social bias in AI-generated images. We test 13 commercially available LMMs using 75 procedurally-generated, gender-neutral prompts to generate people in stereotypically-male, stereotypically-female, and non-stereotypical professions. We then use a validated LLM-as-a-judge system to score the 965 resulting images for gender representation. Our results reveal (p < .001 for all): 1) LMMs systematically not only reproduce but actually amplify occupational gender stereotypes relative to real-world labor data, generating men in 93.0% of images for male-stereotyped professions but only 22.5% for female-stereotyped professions; 2) Models exhibit a strong default-male bias, generating men in 68.3% of the time for non-stereotyped professions; and 3) The extent of bias varies dramatically across models, with overall male representation ranging from 46.7% to 73.3%. Notably, the top-performing model de-amplified gender stereotypes and approached gender parity, achieving the highest fairness scores. This variation suggests high bias is not an inevitable outcome but a consequence of design choices. Our work provides the most comprehensive cross-model benchmark of gender bias to date and underscores the necessity of standardized, automated evaluation tools for promoting accountability and fairness in AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。