arXiv:2511.06056cs.CRcs.CV2025-11综述被引 2

系统梳理2020-2025年身份证件伪造攻击检测研究,揭示数据与模型泛化的核心矛盾。

Identity Card Presentation Attack Detection: A Systematic Review

  • 采用PRISMA方法系统分析2020-2025年AI驱动的证件伪造检测进展
  • 发现从CNN到微伪造分析、再到大模型的显著技术演进路径
  • 指出真实数据与合成数据间的'现实差距'和'合成效用差距'问题

远程身份验证是现代数字安全的关键环节,但极易受到利用伪造或篡改证件的呈现攻击(PAs)威胁。尽管深度学习(DL)推动了呈现攻击检测(PAD)的发展,该领域仍受限于数据稀缺及模型在不同证件类型和新型攻击方法下的泛化能力不足。本文依据PRISMA方法开展系统文献回顾(SLR),全面分析2020至2025年间基于人工智能的证件类PAD研究现状。分析显示方法学发生显著演进:由标准卷积神经网络(CNNs)转向专门的微伪造特征分析,并近期逐步采纳大规模基础模型(FMs),标志着领域重大转变。我们识别出制约进展的核心悖论——模型在大规模私有数据集上验证与在有限公共数据集(通常为模拟或合成数据)评估之间存在显著“现实差距”,导致研究结果难以复现。此外,还揭示“合成效用差距”:合成数据生成虽是应对数据短缺的主要学术回应,却常无法准确预测其在实际伪造检测中的有效性,易致模型过拟合于生成伪影而非真实攻击特征。本综述整合研究成果,识别关键空白,并提供明确参考框架,为构建安全、鲁棒且具备全球泛化能力的PAD系统指明未来研究路径。

原文摘要 · Abstract (English)

Remote identity verification is essential for modern digital security; however, it remains highly vulnerable to sophisticated Presentation Attacks (PAs) that utilise forged or manipulated identity documents. Although Deep Learning (DL) has driven advances in Presentation Attack Detection (PAD), the field is fundamentally limited by a lack of data and the poor generalisation of models across various document types and new attack methods. This article presents a systematic literature review (SLR) conducted in accordance with the PRISMA methodology, aiming to analyse and synthesise the current state of AI-based PAD for identity documents from 2020 to 2025 comprehensively. Our analysis reveals a significant methodological evolution: a transition from standard Convolutional Neural Networks (CNNs) to specialised forensic micro-artefact analysis, and more recently, the adoption of large-scale Foundation Models (FMs), marking a substantial shift in the field. We identify a central paradox that hinders progress: a critical "Reality Gap" exists between models validated on extensive, private datasets and those assessed using limited public datasets, which typically consist of mock-ups or synthetic data. This gap limits the reproducibility of research results. Additionally, we highlight a "Synthetic Utility Gap," where synthetic data generation the primary academic response to data scarcity often fails to predict forensic utility. This can lead to model overfitting to generation artefacts instead of the actual attack. This review consolidates our findings, identifies critical research gaps, and provides a definitive reference framework that outlines a prescriptive roadmap for future research aimed at developing secure, robust, and globally generalizable PAD systems.

身份验证伪造检测深度学习数据缺口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。