用视觉大模型提升人脸伪造检测精度,安全等级下误检率降为2.17%。
DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection

- 以大模型嵌入代替传统特征,捕捉伪造与真实图像差异
- 仅微调少量参数,在高安全场景下误检率降至2.17%
- 适合边境管控等对安全性要求极高的实际部署场景
本文提出DifFoundMAD,一种参数高效的差分形态攻击检测框架,利用视觉基础模型(FM)的泛化能力,捕捉可疑伪造图像与真实活体图像之间的差异。不同于依赖人脸识别嵌入或手工特征的传统D-MAD系统,DifFoundMAD沿用标准差分范式,但将底层表示空间替换为从基础模型提取的嵌入。通过轻量级微调与类别平衡优化,仅更新少量参数,同时保留基础模型丰富的表征先验。在标准D-MAD基准上的跨数据库评估表明,该方法持续优于现有最先进系统,尤其在边境控制等高安全要求场景中表现显著:当前最优系统在高安全级别下的误检率由6.16%降至2.17%。
原文摘要 · Abstract (English)
In this work, we introduce DifFoundMAD, a parameter-efficient D-MAD framework that exploits the generalisation capabilities of vision foundation models (FM) to capture discrepancies between suspected morphs and live capture images. In contrast to conventional D-MAD systems that rely on face recognition embeddings or handcrafted feature differences, DifFoundMAD follows the standard differential paradigm while replacing the underlying representation space with embeddings extracted from FMs. By combining lightweight finetuning with class-balanced optimisation, the proposed method updates only a small subset of parameters while preserving the rich representational priors of the underlying FMs. Extensive cross-database evaluations on standard D-MAD benchmarks demonstrate that DifFoundMAD achieves consistent improvements over state-of-the-art systems, particularly at the strict security levels required in operational deployments such as border control: The error rates reported in the current state-of-the-art were reduced from 6.16% to 2.17% for high-security levels using DifFoundMAD
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。