arXiv:2505.15332cs.CV2025-05中稿 · IEEE International…被引 4

用多模态大模型实现零样本差异人脸伪造攻击检测,提升准确率与可解释性。

Towards Zero-Shot Differential Morphing Attack Detection with Multimodal Large Language Models

  • 设计思维链提示,提升大模型推理可靠性与决策透明度。
  • 在54人护照录入数据上验证,ChatGPT-4o对GAN类伪造检测更优。
  • 对比两种模型表现,适合关注可解释性的安全系统研发者。

利用多模态大语言模型(LLMs)为提升真实生物识别场景中差异人脸伪造攻击检测(D-MAD)的准确性与可解释性提供了新思路。本文首次基于真实受试者数据,将多模态LLMs应用于D-MAD。为有效利用模型,设计基于思维链(CoT)的提示工程,降低拒答率并增强决策可解释性。贡献包括:(1) 首次在真实数据上应用多模态LLMs进行D-MAD;(2) CoT提示工程提升响应可靠性与可解释性;(3) 在54名个体的护照注册数据上完成定性和定量基准测试;(4) 对比ChatGPT-4o与Gemini在检测准确率与决策透明性上的表现。实验表明,ChatGPT-4o在应对基于GAN的伪造攻击时优于Gemini,但两者在复杂条件下均表现不佳;Gemini解释更一致,而ChatGPT-4o更具鲁棒性但拒答率更高。

原文摘要 · Abstract (English)

Leveraging the power of multimodal large language models (LLMs) offers a promising approach to enhancing the accuracy and interpretability of morphing attack detection (MAD), especially in real-world biometric applications. This work introduces the use of LLMs for differential morphing attack detection (D-MAD). To the best of our knowledge, this is the first study to employ multimodal LLMs to D-MAD using real biometric data. To effectively utilize these models, we design Chain-of-Thought (CoT)-based prompts to reduce failure-to-answer rates and enhance the reasoning behind decisions. Our contributions include: (1) the first application of multimodal LLMs for D-MAD using real data subjects, (2) CoT-based prompt engineering to improve response reliability and explainability, (3) comprehensive qualitative and quantitative benchmarking of LLM performance using data from 54 individuals captured in passport enrollment scenarios, and (4) comparative analysis of two multimodal LLMs: ChatGPT-4o and Gemini providing insights into their morphing attack detection accuracy and decision transparency. Experimental results show that ChatGPT-4o outperforms Gemini in detection accuracy, especially against GAN-based morphs, though both models struggle under challenging conditions. While Gemini offers more consistent explanations, ChatGPT-4o is more resilient but prone to a higher failure-to-answer rate.

攻击检测多模态模型零样本可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。