无需训练即可检测人脸伪造,开源多模态模型表现超预期
Emergent Morphing Attack Detection in Open Multi-modal Large Language Models
- 直接用预训练模型识别单张图像中的伪造痕迹
- 最佳模型在等错误率上比专用系统低23%以上
- 适合安全验证与可解释性要求高的生物识别场景
人脸拼接攻击威胁生物特征验证安全,但现有检测系统需针对性训练且泛化能力差。本文首次对开源多模态大语言模型(MLLM)进行零样本单图检测评估,采用公开权重与标准化协议。实验表明,多个MLLM在不经过微调或领域适配的情况下,即具备非平凡的判别能力;其中LLaVA1.6-Mistral-7B模型达到当前最优性能,等错误率(EER)优于高度优化的专用检测基线至少23%。结果表明,多模态预训练能隐式编码面部细微不一致,实现零样本取证敏感性。该能力为生物识别安全与可解释图像分析提供可靠基础,也为后续通过轻量微调进一步提升精度效率提供了新路径。所有代码与评估协议将在发表后公开。
原文摘要 · Abstract (English)
Face morphing attacks threaten biometric verification, yet most morphing attack detection (MAD) systems require task-specific training and generalize poorly to unseen attack types. Meanwhile, open-source multimodal large language models (MLLMs) have demonstrated strong visual-linguistic reasoning, but their potential in biometric forensics remains underexplored. In this paper, we present the first systematic zero-shot evaluation of open-source MLLMs for single-image MAD, using publicly available weights and a standardized, reproducible protocol. Across diverse morphing techniques, many MLLMs show non-trivial discriminative ability without any fine-tuning or domain adaptation, and LLaVA1.6-Mistral-7B achieves state-of-the-art performance, surpassing highly competitive task-specific MAD baselines by at least 23% in terms of equal error rate (EER). The results indicate that multimodal pretraining can implicitly encode fine-grained facial inconsistencies indicative of morphing artifacts, enabling zero-shot forensic sensitivity. Our findings position open-source MLLMs as reproducible, interpretable, and competitive foundations for biometric security and forensic image analysis. This emergent capability also highlights new opportunities to develop state-of-the-art MAD systems through targeted fine-tuning or lightweight adaptation, further improving accuracy and efficiency while preserving interpretability. To support future research, all code and evaluation protocols will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。