arXiv:2601.09896cs.HCcs.AI2026-01被引 3

审计发现图像审美模型偏袒西方男性视角,强化文化不公。

The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor

  • 通过三组数据审计,发现模型偏好含女性描述的图像,排斥男性与LGBTQ+相关图像
  • 对33万张艺术图像评分显示,西方与日本写实风景、城市和人像获最高分
  • 揭示模型训练数据源集中于英语摄影师和西方AI爱好者,反映深层文化偏见

视觉生成AI模型普遍采用统一的审美标准进行训练,但审美本身与个人品味和文化价值密不可分。本文研究了广泛用于数据集筛选与生成图像质量评估的LAION-Aesthetics Predictor(LAP)模型。通过对三个数据集的审计发现,使用LAP筛选的约12亿张图像中,含女性描述的图像被显著偏好,而含男性或LGBTQ+描述的图像则被系统性过滤。在对约33万张艺术图像的评分中,西方与日本艺术家的写实风景、城市和人物肖像得分最高。数字民族志分析表明,模型训练所用的审美评分主要来自英语摄影师和西方AI爱好者,反映出其背后的文化偏见。该算法的凝视延续了西方艺术史中的帝国主义与男性凝视。研究呼吁开发者摒弃单一审美标准,转向更具包容性的评价体系。

原文摘要 · Abstract (English)

Visual generative AI models are trained using a one-size-fits-all measure of aesthetic appeal. However, what is deemed "aesthetic" is inextricably linked to personal taste and cultural values, raising the question of whose taste is represented in visual generative AI models. In this work, we study an aesthetic evaluation model--LAION-Aesthetics Predictor (LAP)--that is widely used to curate datasets to train visual generative image models, like Stable Diffusion, and evaluate the quality of AI-generated images. To understand what LAP measures, we audited the model across three datasets. First, we examined the impact of aesthetic filtering on the LAION-Aesthetics Dataset (approximately 1.2B images), which was curated from LAION-5B using LAP. We find that the LAP disproportionally filters in images with captions mentioning women, while filtering out images with captions mentioning men or LGBTQ+ people. Then, we used LAP to score approximately 330k images across two art datasets, finding the model rates realistic images of landscapes, cityscapes, and portraits from western and Japanese artists most highly. In doing so, the algorithmic gaze of this aesthetic evaluation model reinforces the imperial and male gazes found within western art history. In order to understand where these biases may have originated, we performed a digital ethnography of public materials related to the creation of LAP. We find that the development of LAP reflects the biases we found in our audits, such as the aesthetic scores used to train LAP primarily coming from English-speaking photographers and western AI-enthusiasts. In response, we discuss how aesthetic evaluation can perpetuate representational harms and call on AI developers to shift away from prescriptive measures of "aesthetics" toward more pluralistic evaluation.

审美评估算法偏见生成模型文化批判

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。