统一检测真假人脸攻击,融合空间与频域特征提升识别精度
FA^{3}-CLIP: Frequency-Aware Cues Fusion and Attack-Agnostic Prompt Learning for Unified Face Attack Detection
- 通过空间与频域特征融合,生成通用活体与伪造提示
- 在多个数据集上达到当前最佳性能,显著提升检测效果
- 适合需要同时应对物理和数字攻击的安防系统应用
人脸识别系统易受物理(如打印照片)和数字(如DeepFake)攻击。现有方法难以同时检测两类攻击,主要因攻击间类内差异大,且仅依赖空间信息无法全面捕捉真实与伪造线索。为此,我们提出统一攻击检测模型FA³-CLIP,引入攻击无关提示学习,融合空间与频域特征,提取通用活体与伪造表征,实现对活体及各类攻击的统一检测。具体而言,语言分支中的攻击无关提示模块生成通用活体与伪造提示,从真实与伪造人脸中提取对应表征,引导模型学习统一特征空间。视觉分支采用双流线索融合框架,利用频域信息补充空间域难捕捉的细微线索;频域流中加入频域压缩块,减少冗余同时保留关键线索多样性。此外,构建新挑战性评估协议以促进统一攻击检测研究。实验表明,该方法在物理与数字攻击检测上均显著优于现有方法,达到当前最优性能。
原文摘要 · Abstract (English)
Facial recognition systems are vulnerable to physical (e.g., printed photos) and digital (e.g., DeepFake) face attacks. Existing methods struggle to simultaneously detect physical and digital attacks due to: 1) significant intra-class variations between these attack types, and 2) the inadequacy of spatial information alone to comprehensively capture live and fake cues. To address these issues, we propose a unified attack detection model termed Frequency-Aware and Attack-Agnostic CLIP (FA\textsuperscript{3}-CLIP), which introduces attack-agnostic prompt learning to express generic live and fake cues derived from the fusion of spatial and frequency features, enabling unified detection of live faces and all categories of attacks. Specifically, the attack-agnostic prompt module generates generic live and fake prompts within the language branch to extract corresponding generic representations from both live and fake faces, guiding the model to learn a unified feature space for unified attack detection. Meanwhile, the module adaptively generates the live/fake conditional bias from the original spatial and frequency information to optimize the generic prompts accordingly, reducing the impact of intra-class variations. We further propose a dual-stream cues fusion framework in the vision branch, which leverages frequency information to complement subtle cues that are difficult to capture in the spatial domain. In addition, a frequency compression block is utilized in the frequency stream, which reduces redundancy in frequency features while preserving the diversity of crucial cues. We also establish new challenging protocols to facilitate unified face attack detection effectiveness. Experimental results demonstrate that the proposed method significantly improves performance in detecting physical and digital face attacks, achieving state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。