顶级AI视觉模型识破假人脸胜过年轻人,但判断力不如人平衡。
Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration
- 用19个视觉语言模型测试假人脸识别能力,方法与人类实验一致。
- gpt-5.6-sol模型达92.1%准确率,超过20-30岁人类(88.5%)。
- 模型过激或保守,缺乏人类的判断平衡,适合安全检测场景。
当前AI图像生成已能制作难以分辨的真人肖像。本文对19个视觉语言模型(VLMs)在相同198张肖像上进行评测——包含真实照片及与之身份匹配的ChatGPT-4o和Imagen 3生成版本,任务设置与此前对1667名成人(总体准确率85%;年龄越大准确率越低)的研究一致。2026年6月发布的14个模型仅达到20-30岁成年人水平。四周后性能突破:5个7月发布的模型中,gpt-5.6-sol实现92.8%平衡准确率(五次重复平均92.1%),显著高于同龄人(88.5%);claude-fable-5则100%检出所有AI图像,平均准确率达91.9%。模型敏感性显著超越年轻人(d'最高达3.4,对比约2.4)。但人类的校准能力仍未被超越:模型判别标准范围从c = -1.10至+1.45,而人类各年龄段均接近零;两位领先模型分别存在+0.44和-0.97偏差,仅有少数中等表现模型接近人类平衡。更换标注样本仍导致约四分之一判断结果反转。
原文摘要 · Abstract (English)
AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。