arXiv:2507.15285cs.CV2025-07被引 2

用视觉语言模型实现无需训练的攻击检测,提升人脸识别系统安全性。

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems

  • 基于上下文学习,直接用预训练模型识别物理和数字攻击
  • 在公开数据集上表现优于部分传统CNN模型,且无需额外训练
  • 适合需要快速部署、缺乏标注数据的安全场景

近期生物特征系统在防范欺诈行为方面取得显著进展。然而,随着检测方法的进步,攻击手段也日益复杂。针对人脸识别系统的攻击可分为物理攻击和数字攻击两类。传统深度学习模型虽在特定场景下表现优异,但在面对新型攻击或环境变化时泛化能力不足。这类模型通常需要大量标注数据训练,而生物特征数据收集面临隐私担忧及实际场景采集困难。本文研究了视觉语言模型(VLM)在该领域的应用,提出一种基于上下文学习的攻击检测框架,用于识别物理呈现攻击与数字变形攻击。聚焦开源模型,首次建立了系统性评估VLM在安全关键场景下通过上下文学习进行定量分析的框架。在多个公开数据库上的实验表明,所提方法在无需资源密集型训练的前提下,对物理与数字攻击检测性能具有竞争力,验证了其在提升攻击检测泛化能力方面的潜力。

原文摘要 · Abstract (English)

Recent advances in biometric systems have significantly improved the detection and prevention of fraudulent activities. However, as detection methods improve, attack techniques become increasingly sophisticated. Attacks on face recognition systems can be broadly divided into physical and digital approaches. Traditionally, deep learning models have been the primary defence against such attacks. While these models perform exceptionally well in scenarios for which they have been trained, they often struggle to adapt to different types of attacks or varying environmental conditions. These subsystems require substantial amounts of training data to achieve reliable performance, yet biometric data collection faces significant challenges, including privacy concerns and the logistical difficulties of capturing diverse attack scenarios under controlled conditions. This work investigates the application of Vision Language Models (VLM) and proposes an in-context learning framework for detecting physical presentation attacks and digital morphing attacks in biometric systems. Focusing on open-source models, the first systematic framework for the quantitative evaluation of VLMs in security-critical scenarios through in-context learning techniques is established. The experimental evaluation conducted on freely available databases demonstrates that the proposed subsystem achieves competitive performance for physical and digital attack detection, outperforming some of the traditional CNNs without resource-intensive training. The experimental results validate the proposed framework as a promising tool for improving generalisation in attack detection.

视觉语言模型攻击检测上下文学习人脸识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。