arXiv:2503.10937cs.CVcs.AI2025-03

用零样本学习让大模型识别换脸攻击,无需训练就能通用且能解释。

ChatGPT Encounters Morphing Attack Detection: Zero-Shot MAD with Multi-Modal Large Language Models and General Vision Models

  • 利用大模型零样本特性,不依赖标注数据直接检测换脸攻击。
  • 在未见过的换脸算法上仍达到高准确率,证明方法泛化性强。
  • 支持生成解释性反馈,适合边境安检等需透明决策的场景。

人脸识别系统日益面临换脸攻击威胁,催生了换脸攻击检测(MAD)算法的发展。然而,现有MAD方法普遍存在对未见数据泛化能力差、缺乏可解释性的问题,限制其在注册站和自动边境管控等实际场景的应用。鉴于多数现有MAD算法依赖监督学习,本文探索了一种基于大语言模型(LLM)的零样本学习新范式。提出两类零样本MAD方法:一类使用通用视觉模型,通过计算独立支持集的均值嵌入实现检测,无需使用被篡改图像;另一类则利用最先进的GPT-4 Turbo API,结合精心设计的提示词完成任务。为验证零样本MAD的可行性及方法有效性,构建了一个包含多种未见换脸算法的打印-扫描换脸数据集,模拟真实复杂场景。实验结果表明,该方法具备显著检测准确率,验证了零样本学习在MAD任务中的适用性。此外,研究发现多模态大模型如ChatGPT在未训练过的MAD任务上表现出优异泛化能力,并具备生成解释与指导的独特优势,有助于提升实际应用中的透明度与可用性。

原文摘要 · Abstract (English)

Face Recognition Systems (FRS) are increasingly vulnerable to face-morphing attacks, prompting the development of Morphing Attack Detection (MAD) algorithms. However, a key challenge in MAD lies in its limited generalizability to unseen data and its lack of explainability-critical for practical application environments such as enrolment stations and automated border control systems. Recognizing that most existing MAD algorithms rely on supervised learning paradigms, this work explores a novel approach to MAD using zero-shot learning leveraged on Large Language Models (LLMs). We propose two types of zero-shot MAD algorithms: one leveraging general vision models and the other utilizing multimodal LLMs. For general vision models, we address the MAD task by computing the mean support embedding of an independent support set without using morphed images. For the LLM-based approach, we employ the state-of-the-art GPT-4 Turbo API with carefully crafted prompts. To evaluate the feasibility of zero-shot MAD and the effectiveness of the proposed methods, we constructed a print-scan morph dataset featuring various unseen morphing algorithms, simulating challenging real-world application scenarios. Experimental results demonstrated notable detection accuracy, validating the applicability of zero-shot learning for MAD tasks. Additionally, our investigation into LLM-based MAD revealed that multimodal LLMs, such as ChatGPT, exhibit remarkable generalizability to untrained MAD tasks. Furthermore, they possess a unique ability to provide explanations and guidance, which can enhance transparency and usability for end-users in practical applications.

换脸检测零样本学习大模型应用可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。