arXiv:2602.20569cs.CV2026-02被引 4

首个专用于检测AI伪造金融单据的基准,揭示现有工具基本失效。

AIForge-Doc: A Benchmark for Detecting AI-Forged Tampering in Financial and Form Documents

  • 用AI图像修复工具伪造真实票据和表单中的数字字段,生成高精度篡改图
  • 4061张伪造图像覆盖9种语言,标注了像素级篡改区域,适配DocTamper格式
  • 三类主流检测器性能暴跌,证明当前AI伪造已突破现有识别能力

我们提出AIForge-Doc,首个专注于扩散模型驱动的图像修复(inpainting)在金融与表单文档中伪造行为的基准数据集,具备像素级标注。现有伪造文档数据集依赖传统编辑工具(如Adobe Photoshop、GIMP),导致前沿检测方法对日益增长的AI伪造威胁完全失效。AIForge-Doc通过使用Gemini 2.5 Flash Image与Ideogram v2 Edit两个AI修复API,在四个公开文档数据集(CORD、WildReceipt、SROIE、XFUND)上系统性地篡改真实票据与表单中的数值字段,共生成4,061张伪造图像,涵盖九种语言,并以DocTamper兼容格式提供像素级篡改区域掩码。我们评估了三种代表性检测器——TruFor、DocTamper及零样本GPT-4o判断器,发现所有方法性能严重下降:TruFor在零样本、跨分布情况下AUC=0.751(原NIST16上为0.96);DocTamper在分布内测试中AUC=0.563(原为0.98),像素级交并比IoU=0.020;GPT-4o仅达0.509,几乎处于随机水平。结果表明,当前自动化检测器与视觉语言模型已无法区分此类AI伪造内容,凸显该问题为文档取证领域的新挑战。

原文摘要 · Abstract (English)

We present AIForge-Doc, the first dedicated benchmark targeting exclusively diffusion-model-based inpainting in financial and form documents with pixel-level annotation. Existing document forgery datasets rely on traditional digital editing tools (e.g., Adobe Photoshop, GIMP), creating a critical gap: state-of-the-art detectors are blind to the rapidly growing threat of AI-forged document fraud. AIForge-Doc addresses this gap by systematically forging numeric fields in real-world receipt and form images using two AI inpainting APIs -- Gemini 2.5 Flash Image and Ideogram v2 Edit -- yielding 4,061 forged images from four public document datasets (CORD, WildReceipt, SROIE, XFUND) across nine languages, annotated with pixel-precise tampered-region masks in DocTamper-compatible format. We benchmark three representative detectors -- TruFor, DocTamper, and a zero-shot GPT-4o judge -- and find that all existing methods degrade substantially: TruFor achieves AUC=0.751 (zero-shot, out-of-distribution) vs. AUC=0.96 on NIST16; DocTamper achieves AUC=0.563 vs. AUC=0.98 in-distribution, with pixel-level IoU=0.020; GPT-4o achieves only 0.509 -- essentially at chance -- confirming that AI-forged values are indistinguishable to automated detectors and VLMs. These results demonstrate that AIForge-Doc represents a qualitatively new and unsolved challenge for document forensics.

文档伪造AI检测扩散模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。