人类识脸靠的是推断面部成因、忽略干扰因素的深层机制。
Human face perception reflects inverse-generative and naturalistic discriminative objectives

- 用反向渲染等任务训练的模型更贴近人脑识脸方式
- 自然图像训练的模型比合成数据训练的更准确
- 适合研究视觉认知与神经网络对比的学者
人类识脸的表征机制仍是计算谜题。深度神经网络为这一现象提供了机制假说,但理论不同的模型在随机人脸样本上预测结果常难以区分。为揭示这些假说的诊断差异,我们比较了六个共享架构但训练任务不同的神经网络模型,使用能引发模型分歧的“争议性”人脸对,以及随机采样的人脸对。通过864名被试在不同真实度和姿态变化下的脸貌差异判断测试,发现优先关注高层不变结构的模型(如反向渲染、人脸识别或物体分类训练)最符合人类判断。此外,自然图像训练的模型普遍优于合成数据训练的模型。结果表明,人类识脸受制于推断面部外观潜在成因、抑制无关变化、并由自然图像统计特性调优的机制。
原文摘要 · Abstract (English)
The perceptual representations supporting our ability to recognize faces remain a computational mystery. Deep neural networks offer mechanistic hypotheses for human face perception, but theoretically distinct models often make indistinguishable representational predictions for randomly sampled faces. To expose diagnostic differences among these hypotheses, we compared six neural network models sharing an architecture but trained on distinct tasks, using face pairs optimized to elicit contrasting model predictions ("controversial" pairs) alongside randomly sampled pairs. We tested model predictions against face-dissimilarity judgments from 864 human participants across stimulus sets differing in realism and pose variation. Models prioritizing high-level, invariant structures (trained via inverse rendering, face identification, or object classification) most robustly matched human judgments. Furthermore, models trained on natural images typically outperformed synthetic-trained counterparts. Together, these findings suggest that human face perception is shaped by mechanisms that infer latent causes of facial appearance, discount nuisance variation, and are tuned by natural image statistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。