arXiv:2603.16629cs.CVcs.AI2026-03中稿 · 14th International…被引 2

用大模型生成人脸比对解释,发现其常凭空编造特征。

MLLM-based Textual Explanations for Face Comparison

  • 结合图像与传统识别得分,提升比对准确率
  • 即使判断正确,解释中70%依赖无证据的虚构特征
  • 提出新评估框架,量化解释的可信度证据

多模态大语言模型(MLLM)被用于生成人脸识别决策的自然语言解释。尽管这提升了可解释性,但其在非受限人脸图像上的可靠性尚未充分探索。本文系统分析了在具有挑战性的IJB-S数据集上,面对极端姿态变化和监控影像时,MLLM生成的解释表现。结果显示,即便模型做出正确验证判断,其伴随的解释仍频繁依赖无法验证或虚构的人脸属性,这些属性缺乏视觉证据支持。我们进一步研究将传统人脸识别系统输出的评分与决策结果与输入图像一同提供给MLLM的影响。虽然该方法提升了分类验证性能,但并未确保解释的忠实性。为此,我们引入基于似然比的评估框架,以衡量文本解释的证据强度。研究揭示了当前MLLM在可解释人脸识别中的根本局限,并强调了在生物识别应用中建立可靠可信解释评价体系的必要性。代码已公开于https://github.com/redwankarimsony/LR-MLLMFR-Explainability。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have recently been proposed as a means to generate natural-language explanations for face recognition decisions. While such explanations facilitate human interpretability, their reliability on unconstrained face images remains underexplored. In this work, we systematically analyze MLLM-generated explanations for the unconstrained face verification task on the challenging IJB-S dataset, with a particular focus on extreme pose variation and surveillance imagery. Our results show that even when MLLMs produce correct verification decisions, the accompanying explanations frequently rely on non-verifiable or hallucinated facial attributes that are not supported by visual evidence. We further study the effect of incorporating information from traditional face recognition systems, viz., scores and decisions, alongside the input images. Although such information improves categorical verification performance, it does not consistently lead to faithful explanations. To evaluate the explanations beyond decision accuracy, we introduce a likelihood-ratio-based framework that measures the evidential strength of textual explanations. Our findings highlight fundamental limitations of current MLLMs for explainable face recognition and underscore the need for a principled evaluation of reliable and trustworthy explanations in biometric applications. Code is available at https://github.com/redwankarimsony/LR-MLLMFR-Explainability.

可解释性人脸比对大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。