arXiv:2609.02512cs.CVcs.HC2026-09

AI模型普遍高估人脸魅力,但能准确排序。

Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness

  • 对比2513人与4大AI模型评分,发现AI更偏爱人脸且范围狭窄
  • AI与人类评分相关性强,可准确反映颜值排名顺序
  • 仅年龄是共性判断依据,其他特征表现不一,适合研究审美偏差

多模态大语言模型(MLLMs)在美容、商业和美学领域的吸引力评估日益流行。本研究通过预注册探索性实验,将2513名参与者的人类评分与Claude、Gemini、GPT和Grok四款主流商用AI模型的评分进行比较。结果显示,MLLMs系统性地对人脸给予更高评价,且评分范围更窄,无法在绝对值上复现人类判断。然而,它们与人类评分具有强相关性,能准确追踪人脸的相对排序。不同模型可能依赖不同线索:仅有面部年龄是人类与所有模型共同的预测因子,而种族和性别则表现出不一致模式。各模型间一致性高,除Grok外,其与人类评分的一致性也最低。结论表明,当前商用MLLMs虽能近似人类颜值排序,但普遍存在对人脸魅力的系统性高估。

原文摘要 · Abstract (English)

Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular amongst users, companies, and aestheticians. This raises the question of whether these AI models can accurately reflect human judgments of attractiveness. In a pre- registered exploratory study, we compared the attractiveness ratings of 2,513 human participants to four widely used commercial AI models: Claude, Gemini, GPT, and Grok. Results showed that MLLMs systematically rate faces more favourably and within a narrower range than humans and, at the time of study, do not reproduce human ratings in absolute terms. However, MLLMs exhibit strong correlations with human attractiveness judgments, accurately tracking the rank-ordering of faces. MLLMs may judge faces by different cues than humans; only face age was a predictor of facial attractiveness in both humans and MLLMs, with inconsistent patterns across models for ethnicity and gender. AI models strongly agree with one another, except for Grok, which also showed the lowest agreement with humans. Our findings suggest that while they may be able to approximate rank-orderings of human attractiveness, current off-the-shelf commercial MLLMs systematically overrate the beauty of human faces.

AI审美人脸评估多模态模型偏见分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。