arXiv:2608.24430cs.CV2026-08

用多个视觉语言模型融合提升人脸识别可解释性与准确率。

Vision Language Model Fusion for Explainable Face Recognition

论文配图:Vision Language Model Fusion for Explainable Face Recognition
图 1 · 摘自论文原文
  • 多模型联合决策,融合分数、文本解释和图像
  • 融合模型准确率超越现有最优VLM和专用模型
  • 生成更丰富可信的解释,适合需要透明验证的场景

负责任地部署人脸识别系统不仅要求高精度,还需提供可解释、可审计的证据,使用户能理解、评估并质疑系统决策。视觉语言模型(VLMs)通过结合视觉分析与自然语言推理,为可解释人脸识别提供了良好基础。然而单一模型可能限制准确率与解释质量。本文评估了四种VLM作为独立识别系统的表现,并提出一种融合框架:两个源模型提供相似度分数与文本解释,第三个模型作为决策者。研究比较了四种融合场景,逐步向决策模型输入分数、解释、人脸图像及组合模态。结果表明,多VLM融合的价值不仅在于提升识别性能,还能提供互补的解释视角,使决策解释更丰富,增强透明性、可审计性与错误分析能力。所提多模态融合模型在结合决策分数、解释与人脸图像后,准确率超过当前最先进VLM与领域专用模型,且生成的融合解释预计比单个模型更稳健。

原文摘要 · Abstract (English)

Responsible deployment of face verification systems requires more than accurate decisions: systems should also provide interpretable and auditable evidence that enables users to understand, assess, and challenge their decisions. Vision-language models (VLMs) provide a promising foundation for explainable face recognition by combining visual analysis with natural-language reasoning. However, relying on a single model may further limit the decision accuracy as well as provided explanations. This work therefore investigates whether multiple VLMs can be combined to improve recognition accuracy, and to enrich the explanations associated with those decisions. This work evaluates four VLMs as standalone face verification systems and subsequently proposes a fusion framework, where two source models provide similarity scores and textual justifications and a third VLM acts as a decider model. Four different fusion scenarios are considered, progressively providing the decider model with scores, justifications, face images, and combinations of these modalities. Overall, the findings suggest that the value of multi-VLM fusion extends beyond recognition performance. VLMs can provide complementary justifications and perspectives that enable richer explanations of face recognition decisions, supporting greater transparency, auditability, and error analysis. This is relevant to the development of responsible explainable face verification systems, where users and operators should be able to understand not only the final decision but also the evidence and potential sources underlying it. The proposed multimodal VLM, which combines decision scores, explanations, and face images, achieves higher recognition accuracy than state-of-the-art VLMs and domain-specific face recognition models, while also providing fused explanations that are expected to be more robust than those generated by individual VLMs.

可解释AI人脸识别多模型融合视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。