构建真实法庭证据的生成数据集,助力检测AI伪造文书。
The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence

- 用先进生成工具构造多类伪造文书,覆盖局部修改到完全伪造。
- 数据集含多种文档类型和篡改策略,支持严格评估检测模型。
- 专为司法场景设计,适合研究证据真实性验证的学者与工程师。
生成模型日益能制造逼真文件,对司法系统中依赖证据真实性的流程构成直接挑战,如收据、通信记录等。与社交媒体或学术文档不同,司法证据常经细微、局部修改,保持整体可信度但改变法律含义。当前自动化检测进展受限,主要因缺乏适配司法需求的训练与评估数据。现有资源多聚焦人脸或自然景观图像,或仅涵盖特定类型的学术/社交文档,无法反映真实证据的结构、多样性和篡改模式。因此,现有检测系统未必学习到适用于司法场景的有效信号。我们提出CIFAR合成证据语料库(CIFAR Synthetic Evidence Corpus),旨在实现真实且可控条件下证据验证的严谨评估。该语料库涵盖多种文档类别及从字段级编辑到完整伪造的篡改策略,采用多样化先进生成工具构建。其设计可系统性变化篡改复杂度与生成方法,并确保训练与测试数据在来源上分离,以模拟真实世界中的泛化挑战。
原文摘要 · Abstract (English)
The growing ability of generative models to produce realistic documents poses a direct challenge to evidentiary workflows in the justice system and the courts, where decisions increasingly depend on the authenticity of evidence such as receipts, communications, and administrative records. Unlike social media or academic settings, evidentiary documents are often only subtly altered, with small, localized edits that preserve overall plausibility while changing legal meaning. Yet progress on automated detection remains limited, largely due to the absence of suitable training and evaluation data especially suited for the justice system requirements. Existing resources are either focused on photos of human faces or natural scenery or on narrowly scoped academic or social media document types, and do not capture the structure, diversity, or manipulation patterns characteristic of real-world evidentiary data. As a result, current detection systems do not necessarily learn meaningful signals appropriate for the justice system. We introduce the CIFAR Synthetic Evidence Corpus, a dataset designed to enable rigorous evaluation of evidence verification under realistic and controlled conditions. The corpus spans multiple document families and a spectrum of manipulation strategies, from small field-level edits to complete document fabrication, and is constructed using a diverse set of state-of-the-art generative tools. It is organized to systematically vary both manipulation complexity and generation method, while enforcing source-level separation between training and test data to reflect real-world generalization challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。