构建真实感伪造证件检测数据集,助力反欺诈系统研发。
FantasyID: A dataset for detecting digital manipulations of ID-documents
- 构建无法律风险的伪造证件数据集,含真实人脸与多语言设计。
- 现有算法在真实场景下误检率超50%,暴露检测瓶颈。
- 适合安全、金融领域研究者用于评估伪造检测模型性能。
图像生成技术的进步使得恶意用户可轻松创建伪造图像,对广泛使用的KYC应用构成严重威胁,亟需鲁棒的身份证件伪造检测系统。本文提出一个公开可用(含商业用途)的新数据集FantasyID,模拟真实世界身份证件但不涉及合法文件篡改,且不含生成人脸或模板水印。FantasyID包含多种设计风格、语言及真实人物面部的身份证件。为模拟真实KYC场景,这些证件经打印并用三种不同设备拍摄,构成真实样本类。我们使用现有生成工具模拟恶意攻击者可能实施的数字伪造/注入攻击。当前最先进的伪造检测算法(如TruFor、MMFusion、UniFD、FatFormer)在接近实际应用的评估条件下表现不佳:当验证集上设置10%假阳性率时,测试集上假阴性率普遍接近50%。实验表明,FantasyID具备足够复杂性,可作为伪造检测算法的有效评估基准。
原文摘要 · Abstract (English)
Advancements in image generation led to the availability of easy-to-use tools for malicious actors to create forged images. These tools pose a serious threat to the widespread Know Your Customer (KYC) applications, requiring robust systems for detection of the forged Identity Documents (IDs). To facilitate the development of the detection algorithms, in this paper, we propose a novel publicly available (including commercial use) dataset, FantasyID, which mimics real-world IDs but without tampering with legal documents and, compared to previous public datasets, it does not contain generated faces or specimen watermarks. FantasyID contains ID cards with diverse design styles, languages, and faces of real people. To simulate a realistic KYC scenario, the cards from FantasyID were printed and captured with three different devices, constituting the bonafide class. We have emulated digital forgery/injection attacks that could be performed by a malicious actor to tamper the IDs using the existing generative tools. The current state-of-the-art forgery detection algorithms, such as TruFor, MMFusion, UniFD, and FatFormer, are challenged by FantasyID dataset. It especially evident, in the evaluation conditions close to practical, with the operational threshold set on validation set so that false positive rate is at 10%, leading to false negative rates close to 50% across the board on the test set. The evaluation experiments demonstrate that FantasyID dataset is complex enough to be used as an evaluation benchmark for detection algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。