arXiv:2501.01720cs.CV2025-01AAAI被引 22

用大模型让人脸识别防伪可解释,提升跨场景泛化能力

Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models

  • 将防伪任务转为视觉问答,生成自然语言解释
  • 在12个数据集上超越现有方法,跨域准确率提升显著
  • 适合需要可解释性与鲁棒性的安防系统开发者

人脸反欺骗(FAS)对保障人脸识别系统的安全与可靠性至关重要。现有方法多为二分类任务,仅输出置信度而无解释,且在新环境或未见欺骗类型下泛化能力有限。本文提出基于多模态大模型的可解释人脸反欺骗框架I-FAS,将FAS任务转化为可解释的视觉问答(VQA)模式。我们设计了欺骗感知的描述生成与过滤(SCF)策略,为FAS图像生成高质量描述,以自然语言增强模型监督。为缓解训练中噪声描述的影响,提出不对称语言模型(L-LM)损失函数,分离判断与解释的损失计算,优先优化判断部分。此外,设计全局感知连接器(GAC),对齐多层次视觉特征与语言模型。在标准及新构建的One to Eleven跨域基准(共12个公开数据集)上的大量实验表明,该方法显著优于当前最先进方法。

原文摘要 · Abstract (English)

Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tasks, providing confidence scores without interpretation. They exhibit limited generalization in out-of-domain scenarios, such as new environments or unseen spoofing types. In this work, we introduce a multimodal large language model (MLLM) framework for FAS, termed Interpretable Face Anti-Spoofing (I-FAS), which transforms the FAS task into an interpretable visual question answering (VQA) paradigm. Specifically, we propose a Spoof-aware Captioning and Filtering (SCF) strategy to generate high-quality captions for FAS images, enriching the model's supervision with natural language interpretations. To mitigate the impact of noisy captions during training, we develop a Lopsided Language Model (L-LM) loss function that separates loss calculations for judgment and interpretation, prioritizing the optimization of the former. Furthermore, to enhance the model's perception of global visual features, we design a Globally Aware Connector (GAC) to align multi-level visual representations with the language model. Extensive experiments on standard and newly devised One to Eleven cross-domain benchmarks, comprising 12 public datasets, demonstrate that our method significantly outperforms state-of-the-art methods.

人脸反欺骗多模态大模型可解释性跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。