构建首个真实身份证数字伪造数据集,解决隐私与检测难题。
FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking

- 用图像修复技术生成高保真伪造样本,模拟真实攻击
- 通过几何约束提取近520万隐私脱敏补丁,保障合规
- 实测主流模型检测率仅32.45%,凸显防御挑战
身份证件认证依赖复杂高频安全图案的结构完整性。然而,先进生成式AI可注入局部高保真篡改,制造绕过常规验证的欺骗性攻击。训练鲁棒图像取证模型受隐私法规限制,被迫使用缺乏真实视觉特征的合成模板。为此,我们提出FakeIDet3-DB,首个涵盖真实政府颁发身份证的数字篡改综合数据库。该库包含经典(如复制-粘贴)和生成式AI驱动的篡改(如人脸替换、图像修补),并采用先进图像修复技术抑制视觉伪影。为符合严格数据保护法规(如GDPR),我们基于补丁框架设计隐私保护方案。为最大化取证效用,将真实身份证的隐私补丁提取建模为几何约束图像处理问题。提出PACE算法——伪匿名上下文补丁提取,结合积分图映射与距离驱动非极大值抑制(NMS)。该算法高效生成去标识化掩码,防止个人识别信息泄露,同时最大化周边区域语义密度,共从6400+张真实/伪造身份证中提取近520万补丁。进一步评估显示,当前最先进模型在应对生成式与经典技术攻击时均表现不佳:检测错误率(EER)达32.45%,定位准确率(AUC-ROC)为83.48%。
原文摘要 · Abstract (English)
Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Generative AI models can now inject localized, high-fidelity manipulations, creating deceptive attacks that bypass standard verification. Training robust image forensic models to detect these anomalies is hindered by privacy regulations, forcing reliance on synthetic templates lacking the intricate visual patterns of real IDs. To bridge this domain gap, we introduce FakeIDet3-DB, the first comprehensive database of digital manipulations on real, government-issued IDs. FakeIDet3-DB encompasses classical (e.g., copy-move) and Generative AI-driven manipulations (e.g., face-swapping, inpainting) enhanced with advanced image refinement procedures to suppress visual artifacts. In addition, to comply with strict data protection regulations (e.g., GDPR), we adopt a recently-proposed framework based on patches. In order to maximize forensic utility, we formulate privacy-aware patch extraction from a real ID as a geometrically constrained image processing problem. We propose PACE, a Pseudo-Anonymized Contextual patch Extraction algorithm, which leverages Integral Image mapping and distance-driven Non-Maximum Suppression (NMS). PACE efficiently contours anonymization masks that prevent Personally Identifiable Information (PII) leakage while maximizing semantic density in peri-censorship regions, yielding almost 5.2M patches extracted from more than 6.4K images from real/fake IDs. Furthermore, an extensive evaluation of the proposed FakeIDet3-DB is performed using state-of-the-art models, showcasing they all struggle to detect and locate attacks coming from generative and classic techniques (32.45\% EER in detection and 83.48\% AUC-ROC in localization).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。