AI美颜评分模型存在种族偏见,且会放大社会审美不公。
Analysis of Bias in Deep Learning Facial Beauty Regressors
- 对比分析两大数据集训练的模型,发现跨族裔评分差异显著
- 在公平数据集上仍存在显著偏差,仅4.8%-9.5%满足平等标准
- 揭示算法可能固化社会审美偏见,适合关注AI伦理的研究者
即使来自看似均衡的数据源,人工智能系统仍可能引入偏见,而AI面部美学预测尤其存在基于种族的偏差。本文通过在SCUT-FBP5500和MEBeauty数据集上训练的模型进行对比分析,警示人工智能在塑造审美标准中的潜在作用。采用严格的统计验证(Kruskal-Wallis H检验与事后邓恩分析),结果表明,两个模型在公平的FairFace数据集上仍表现出显著的族裔间预测差异(p < 0.001)。跨数据集验证显示,算法不仅未缓解偏见,反而放大了社会性审美偏见,体现在预测与误差的不平等上。研究指出当前AI美学预测方法存在根本缺陷:仅有4.8%-9.5%的族裔间比较满足分布公平性标准。文中提出并深入讨论了相应的缓解策略。
原文摘要 · Abstract (English)
Bias can be introduced to AI systems even from seemingly balanced sources, and AI facial beauty prediction is subject to ethnicity-based bias. This work sounds warnings about AI's role in shaping aesthetic norms while providing potential pathways toward equitable beauty technologies through comparative analysis of models trained on SCUT-FBP5500 and MEBeauty datasets. Employing rigorous statistical validation (Kruskal-Wallis H-tests, post hoc Dunn analyses). It is demonstrated that both models exhibit significant prediction disparities across ethnic groups $(p < 0.001)$, even when evaluated on the balanced FairFace dataset. Cross-dataset validation shows algorithmic amplification of societal beauty biases rather than mitigation based on prediction and error parity. The findings underscore the inadequacy of current AI beauty prediction approaches, with only 4.8-9.5\% of inter-group comparisons satisfying distributional parity criteria. Mitigation strategies are proposed and discussed in detail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。