arXiv:2601.01798cs.CVcs.AI2026-01被引 2

让人脸识别系统会说人话,解释判断依据。

VerLM: Explaining Face Verification Using Natural Language

  • 用视觉语言模型结合双类解释提升决策透明度。
  • 在跨模态迁移中优化视觉差异识别,准确率更高。
  • 适合需要可解释性的安防、金融等高风险场景。

人脸识别系统虽有显著进步,但其决策过程常缺乏透明性。本文提出一种新型视觉语言模型(VLM)用于人脸识别,不仅能准确判断两张人脸是否为同一人,还能明确解释判断依据。模型通过两种互补的解释方式训练:(1)简洁总结影响判断的关键因素;(2)详细描述图像间的具体差异。我们改进了原本用于音频差异分析的先进建模方法,成功将其迁移到视觉输入任务中,显著提升了模型的准确性与可解释性。该模型融合先进的特征提取与推理能力,可清晰表达验证过程。实验表明,该方法优于基线模型和现有模型,展现出视觉语言模型在可解释人脸识别中的巨大潜力。

原文摘要 · Abstract (English)

Face verification systems have seen substantial advancements; however, they often lack transparency in their decision-making processes. In this paper, we introduce an innovative Vision-Language Model (VLM) for Face Verification, which not only accurately determines if two face images depict the same individual but also explicitly explains the rationale behind its decisions. Our model is uniquely trained using two complementary explanation styles: (1) concise explanations that summarize the key factors influencing its decision, and (2) comprehensive explanations detailing the specific differences observed between the images. We adapt and enhance a state-of-the-art modeling approach originally designed for audio-based differentiation to suit visual inputs effectively. This cross-modal transfer significantly improves our model's accuracy and interpretability. The proposed VLM integrates sophisticated feature extraction techniques with advanced reasoning capabilities, enabling clear articulation of its verification process. Our approach demonstrates superior performance, surpassing baseline methods and existing models. These findings highlight the immense potential of vision language models in face verification set up, contributing to more transparent, reliable, and explainable face verification systems.

人脸识别视觉语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。