提出隐私保护的假身份证检测方法与公开数据集
Privacy-Aware Detection of Fake Identity Documents: Methodology, Benchmark, and Improved Algorithms (FakeIDet2)
- 采用图像块分割策略保护身份信息隐私
- 构建超90万张真实/伪造证件图块数据库
- 适合安全、隐私、AI伪造检测研究者使用
互联网应用中的远程用户验证日益重要。常见场景是用户提交身份证照片,平台验证真伪后开放服务。身份证由政府颁发,具有唯一性和不可转让性,但近年来人工智能发展使得伪造真实度极高的物理和合成假证成为可能。现有检测方法面临真实数据匮乏难题,因身份证属敏感信息,个人及机构均不愿共享。本研究贡献包括:1)提出基于图像块的隐私保护检测方法;2)发布新公开数据集FakeIDet2-db,包含从2000张身份证图像中提取的超过90万张真实与伪造证件图块,涵盖多种手机传感器、光照和拍摄角度,并考虑打印、屏幕投射和拼接三类物理攻击;3)提出新型隐私感知的假身份证检测模型FakeIDet2;4)提供可复现的标准基准,涵盖文献中常见的物理与合成攻击类型。
原文摘要 · Abstract (English)
Remote user verification in Internet-based applications is becoming increasingly important nowadays. A popular scenario for it consists of submitting a picture of the user's Identity Document (ID) to a service platform, authenticating its veracity, and then granting access to the requested digital service. An ID is well-suited to verify the identity of an individual, since it is government issued, unique, and nontransferable. However, with recent advances in Artificial Intelligence (AI), attackers can surpass security measures in IDs and create very realistic physical and synthetic fake IDs. Researchers are now trying to develop methods to detect an ever-growing number of these AI-based fakes that are almost indistinguishable from authentic (bona fide) IDs. In this counterattack effort, researchers are faced with an important challenge: the difficulty in using real data to train fake ID detectors. This real data scarcity for research and development is originated by the sensitive nature of these documents, which are usually kept private by the ID owners (the users) and the ID Holders (e.g., government, police, bank, etc.). The main contributions of our study are: 1) We propose and discuss a patch-based methodology to preserve privacy in fake ID detection research. 2) We provide a new public database, FakeIDet2-db, comprising over 900K real/fake ID patches extracted from 2,000 ID images, acquired using different smartphone sensors, illumination and height conditions, etc. In addition, three physical attacks are considered: print, screen, and composite. 3) We present a new privacy-aware fake ID detection method, FakeIDet2. 4) We release a standard reproducible benchmark that considers physical and synthetic attacks from popular databases in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。