arXiv:2603.11442cs.AIcs.CV2026-03被引 3

人类难辨AI伪造账单,因关键错误肉眼看不见。

GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document Forensics

  • 用1235张真假账单对比,测试人与模型识别能力。
  • 人类虽最能察觉差异,但检测准确率低于多个大模型。
  • 核心漏洞是算术错误,仅模型可快速验证。

我们提出 GPT4o-Receipt,一个包含 1,235 张账单图像的基准数据集,每张图对应由 GPT-4o 生成的账单与来自公开数据集的真实账单。该数据集通过五个先进多模态大模型(multimodal LLMs)和由 30 名众包标注者参与的感知研究进行评估。研究发现显著悖论:人类在视觉辨别上表现最强,但二分类检测的 F1 分数却远低于 Claude Sonnet 4 与 Gemini 2.5 Flash。这一矛盾源于核心取证信号为算术错误——肉眼不可见,但可被模型毫秒级验证。此外,五模型评估揭示显著性能差异与校准问题,表明仅靠准确率不足以选择检测器。GPT4o-Receipt、评估框架及全部结果已公开,支持未来人工智能文档溯源研究。

原文摘要 · Abstract (English)

Can humans detect AI-generated financial documents better than machines? We present GPT4o-Receipt, a benchmark of 1,235 receipt images pairing GPT-4o-generated receipts with authentic ones from established datasets, evaluated by five state-of-the-art multimodal LLMs and a 30-annotator crowdsourced perceptual study. Our findings reveal a striking paradox: humans are better at seeing AI artifacts, yet worse at detecting AI documents. Human annotators exhibit the largest visual discrimination gap of any evaluator, yet their binary detection F1 falls well below Claude Sonnet 4 and below Gemini 2.5 Flash. This paradox resolves once the mechanism is understood: the dominant forensic signals in AI-generated receipts are arithmetic errors -- invisible to visual inspection but systematically verifiable by LLMs. Humans cannot perceive that a subtotal is incorrect; LLMs verify it in milliseconds. Beyond the human--LLM comparison, our five-model evaluation reveals dramatic performance disparities and calibration differences that render simple accuracy metrics insufficient for detector selection. GPT4o-Receipt, the evaluation framework, and all results are released publicly to support future research in AI document forensics.

AI伪造文档检测多模态模型算术错误

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。