arXiv:2609.03052cs.CVcs.LG2026-09

生成高仿真身份证件,提升远程身份验证系统评估可靠性

IDSPACE: A Novel Document Generator for Reliable Evaluation of Digital Identity Verification Systems [Extended Technical Report]

论文配图:IDSPACE: A Novel Document Generator for Reliable Evaluation of Digital Identity Verification Systems [Extended Technical Report]
图 1 · 摘自论文原文
  • 用贝叶斯优化自动调节生成参数,提升图像真实度与模型预测一致性
  • 仅需少量真实样本,评估一致性提升15%-45%,训练准确率最高提高9%
  • 支持扫描和手机拍摄文档,适合安全、金融领域研究者使用

随着服务线上化,银行、贷款机构和政府等信任机构需验证远程用户身份。尽管欺诈检测工具广泛可用,但评估与调优仍因身份文件敏感而难以获取真实数据。合成数据生成提供了解决路径,需求明确:我们此前相关工作累计下载超11,000次(来自八个部分)。本文提出IDSpace,从三方面扩展该方向:第一,引入模型引导的贝叶斯优化,仅需目标域少量样本即可优化生成参数,最大化视觉相似性与目标模型预测一致性;第二,将用户设定的元数据(如人口统计、欺诈模式、采集设备)与自动调参项(字体风格、噪声水平、图像质量)解耦,无需低层专业知识即可配置评估;第三,突破模板图像限制,支持扫描件和移动端拍摄的证件。实验表明,相比CycleGAN、扩散修补和非引导优化基线,IDSpace在仅用少量真实样本下,评估一致性提升15%-45%,训练准确率最高提升9%,与目标域的SSIM相似度提升10%。同时发布新数据集,包含359,240张高质量合成证件,覆盖十类欧洲身份证类型。

原文摘要 · Abstract (English)

As services move online, trust institutions such as banks, lenders, and governments must verify the identity of remote users. Fraud detection tools are widely available, but evaluating and fine-tuning them remains difficult because identity documents are sensitive and therefore scarce. Synthetic data generation offers a path forward, and demand is clear: our prior work in this area has been downloaded over $11{,}000$ times (aggregated from eight parts). We introduce IDSpace, extending this line of research in three directions. First, we propose model-guided Bayesian optimization, which tunes generation parameters to maximize both visual similarity and prediction consistency with target-domain models given only a few samples from a target domain. Second, we decouple user-specified metadata (demographics, fraud patterns, capture device) from automatically tuned control parameters (font styles, noise levels, image quality), allowing users to configure evaluations without low-level expertise. Third, we expand beyond template images to support scanned and mobile-captured documents. Experiments show IDSpace improves evaluation consistency by $15-45\%$ over baselines including CycleGAN, diffusion inpainting, and non-guided optimization, using only a few real samples, while improving training accuracy by up to $9\%$ and SSIM similarity with the target domain by $10\%$. We also released a new dataset consisting of $359{,}240$ high-quality synthetic documents across ten European ID types.

身份验证合成数据贝叶斯优化文档生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。