用专业文档不一致检测基准,评估大模型审计能力。
On Finding Inconsistencies in Documents
- 构建人工插入不一致的文档评测集FIND,测试模型审计能力。
- GPT-5识别出64%的人工注入不一致,还发现原稿中136处遗漏问题。
- 适合需要高效文档审核的研究者、法律与金融从业者参考。
学术、法律和金融领域的专业人士需审计文档以避免不一致带来的经济损失、声誉损害和科学错误。语言模型有望大幅提升审计效率。为此,我们提出FIND(Finding INconsistencies in Documents)评测基准,每个样本由领域专家手动插入不一致的长篇技术文档组成。尽管文档复杂,最佳模型GPT-5仍仅能识别64%的插入不一致。令人意外的是,该模型还发现了原始文档中未被察觉的不一致:在50篇arXiv论文中,其提出的196个建议中有136个被判定为真实存在的遗漏问题。然而,即便最优模型仍遗漏近半数不一致,表明不一致检测仍是极具挑战的任务。
原文摘要 · Abstract (English)
Professionals in academia, law, and finance audit their documents because inconsistencies can result in monetary, reputational, and scientific costs. Language models (LMs) have the potential to dramatically speed up this auditing process. To understand their abilities, we introduce a benchmark, FIND (Finding INconsistencies in Documents), where each example is a document with an inconsistency inserted manually by a domain expert. Despite the documents being long, technical, and complex, the best-performing model (gpt-5) recovered 64% of the inserted inconsistencies. Surprisingly, gpt-5 also found undiscovered inconsistencies present in the original documents. For example, on 50 arXiv papers, we judged 136 out of 196 of the model's suggestions to be legitimate inconsistencies missed by the original authors. However, despite these findings, even the best models miss almost half of the inconsistencies in FIND, demonstrating that inconsistency detection is still a challenging task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。