提出可解释的人脸相似度度量,更贴近人类真实感知。
AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations

- 基于认知心理学构建可解释的面部相似度模型
- 在多个人群中显著提升与人类感知的一致性
- 适合用于人脸编辑和隐私保护等场景的可信评估
生成式人脸内容(如人脸编辑、隐私保护)日益影响人们的生活,亟需能忠实反映人类感知的相似度度量。现有方法虽从信号基启发式发展到表示基度量,但仍局限于行为建模,缺乏认知对齐,依赖隐含且虚假的关系,假设存在普适观察者,未考虑不同人群间的内在差异,导致评价失真,误导生成模型调试。本文借鉴认知心理学中关于人脸相似度感知的科学发现:面部特征与配置属性依赖性、非线性心理物理响应、本族偏见。我们构建了FACETS数据集,提出AlignFace——一种可解释、以人为本的脸部相似度度量,通过先验建模编码这些认知原则:使用视觉-语言模型(VLM)编码成对人脸图像与文本属性,门控交叉注意力(CA)提取特定属性的差异表示,概念瓶颈建模(CBM)通过可解释的面部属性约束推理,神经广义加性模型(GAM)建模其非线性影响。实验表明,相比基线度量(包括近期无领域学习的感知度量),AlignFace在多个子群体中显著提升与人类感知的一致性。该工作将学习表征与人类认知过程相连接,为面部图像提供了更透明、对齐的感知评估工具。
原文摘要 · Abstract (English)
Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception. While perceptual evaluation has progressed from signal-based heuristics to representation-based metrics, current approaches are limited to behavioral modeling without cognitive alignment. They rely on implicit and spurious relations while assuming a universal observer, failing to account for inherent variations across diverse human populations. This leads to inaccurate evaluative models of stakeholders and misleading guidance for generative model debugging. Rather than treating perception as a black box, we leverage scientific findings from cognitive psychology of human face similarity perception: dependence on facial featural and configural attributes, nonlinear psychophysical response scaling, and own-group biases. We introduce the FACETS dataset and propose AlignFace, an interpretable, human-aligned, face similarity metric that encodes these cognitive principles through ante-hoc modeling. It employs visual-language modeling (VLM) to encode paired face images and text-based attributes, gated cross-attention (CA) to extract attribute-specific facial difference representations, concept bottleneck modeling (CBM) to constrain reasoning via interpretable face attributes, and neural generalized additive model (GAM) to model their nonlinear influence. Experiments found AlignFace significantly improves alignment with human subpopulation perceptions compared to baseline metrics, including recent domain-free learned perceptual metrics. By bridging learned representations and human cognitive processes, this work enables more transparent and aligned perceptual evaluation metrics for face images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。