arXiv:2508.08939cs.CV2025-08中稿 · ACM Multimedia Wor…被引 2

用多提示聚合提升CLIP零样本识别人脸合成攻击能力

MADPromptS: Unlocking Zero-Shot Morphing Attack Detection with Multiple Prompt Aggregation

  • 设计多个文本提示并聚合其嵌入,增强模型对真实与攻击样本的区分能力
  • 在无任何训练情况下,零样本检测准确率显著高于基线方法
  • 适合安全领域研究者快速部署无需微调的生物特征防御系统

人脸合成攻击检测(MAD)是人脸识别安全中的关键挑战,攻击者可通过插值多人身份信息生成单张伪造人脸图像,使系统误认为该图像属于多个身份。尽管多模态基础模型(如CLIP)具备强大的零样本能力,能联合建模图像与文本,但此前多数研究依赖特定任务的微调,忽视了其直接、通用部署的潜力。本文探索一种纯零样本的MAD方法,仅使用CLIP而无需额外训练或微调,重点在于为每类样本设计并聚合多个文本提示。通过整合多样化提示的嵌入,更好地对齐模型内部表征与MAD任务,捕捉更丰富、多样的真伪样本线索。实验表明,提示聚合显著提升零样本检测性能,验证了通过高效提示工程挖掘基础模型内置多模态知识的有效性。

原文摘要 · Abstract (English)

Face Morphing Attack Detection (MAD) is a critical challenge in face recognition security, where attackers can fool systems by interpolating the identity information of two or more individuals into a single face image, resulting in samples that can be verified as belonging to multiple identities by face recognition systems. While multimodal foundation models (FMs) like CLIP offer strong zero-shot capabilities by jointly modeling images and text, most prior works on FMs for biometric recognition have relied on fine-tuning for specific downstream tasks, neglecting their potential for direct, generalizable deployment. This work explores a pure zero-shot approach to MAD by leveraging CLIP without any additional training or fine-tuning, focusing instead on the design and aggregation of multiple textual prompts per class. By aggregating the embeddings of diverse prompts, we better align the model's internal representations with the MAD task, capturing richer and more varied cues indicative of bona-fide or attack samples. Our results show that prompt aggregation substantially improves zero-shot detection performance, demonstrating the effectiveness of exploiting foundation models' built-in multimodal knowledge through efficient prompt engineering.

零样本检测人脸安全提示工程CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。