揭示人脸数据集中的性别化审美偏见如何通过特征与注意力影响模型行为。
Beyond Performance Disparities: A Three-Level Audit of Representational Harm in CelebA

- 从数据结构、特征权重到注意力分布三层审计,发现性别双重标准被编码其中
- 女性在衰老或男性化特征上遭严重贬损,男性则在年长时被排除评估体系外
- 适合关注公平性、数据偏见与可解释性的研究者阅读
大规模人脸数据集如CelebA广泛应用于计算机视觉,但其标签中的文化偏见仍缺乏深入探讨。公平性研究区分了表征性伤害与分配性伤害,但现有数据集审计多聚焦分类标签,未考察此类伤害如何体现在学习特征与模型注意力中。本文对CelebA进行三层分析:数据结构、学习特征权重与空间注意力,聚焦性别化的衰老与美标准如何在数据中编码并被模型再现。首先,对202,599张图像的分层聚类显示,39个属性形成与文化原型一致的潜在特质组合:表演性女性气质(年轻、化妆、装饰)与职业性男性气质(衰老、胡须、正装)。女性虽整体更常被评为有吸引力,但在衰老或男性化簇中面临严重惩罚。其次,基于XGBoost与SHAP的分析揭示性别特异性效应,如脂肪含量仅降低女性吸引力。第三,Grad-CAM显示,对女性及年轻男性的预测集中于面部中部,而对老年男性的预测则转向发际线和服装等外围线索。老年男性准确率最高,但平均精确率最低,表明其完全被排除在数据集的评价模板之外。文化双重标准由此从媒体表征经由数据标签、特征权重至模型注意力层层传递,造成两类表征性伤害:女性因狭窄评价模板被过度审视,老年男性则被整体排除。仅关注性能差异的公平性度量会掩盖这两种伤害,凸显需在公平性研究中重视表征性伤害。
原文摘要 · Abstract (English)
Large-scale facial datasets like CelebA are widely used in computer vision, yet the cultural biases embedded in their labels remain underexplored. Fairness research has distinguished representational from allocational harms, but audits of computer vision datasets have mostly examined categorical labels, leaving open how such harms appear in learned features and model attention. This paper examines CelebA at three levels: dataset structure, learned feature weights, and spatial attention, focusing on how gendered double standards of ageing and beauty are encoded in the data and reproduced in model behaviour. First, hierarchical clustering of 202,599 images shows that the 39 attributes organise into latent trait bundles aligned with cultural archetypes: performative femininity (youth, makeup, adornment) and professional masculinity (ageing, facial hair, formal attire). Female faces, though more often rated attractive overall, incur steep penalties when assigned to ageing or masculine-coded clusters. Second, XGBoost with SHAP analysis reveal gender-specific effects, such as adiposity reducing attractiveness only for females. Third, Grad-CAM finds that predictions for female and younger male subgroups concentrate on mid-face cues, whereas predictions for older males drift toward peripheral cues such as hair and clothing. Older males attain the highest accuracy but the lowest average precision, indicating categorical exclusion of groups outside the dataset's evaluative templates. Cultural double standards thus pass from media representation into dataset labels, feature weights, and model attention, producing two representational harms: hyper-scrutiny of women under a narrow evaluative template, and exclusion of older men from the scheme entirely. Fairness metrics focused on performance disparities mask both, underscoring the need to address representational harm in fairness research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。