用713组论文错误对测试大模型找科学论文漏洞的能力。
FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
- 用大模型在已审论文中插入能推翻核心结论的错误,构建测试集。
- 顶尖模型GPT 5在10个候选答案中准确找到错误的概率为39.1%。
- 适合研究自动化审稿、模型可信度评估的研究者参考。
错误识别与定位是同行评审的核心任务,但科学产出的爆炸式增长使得专家资源有限的人类评审难以可靠发现错误。近年来大语言模型(LLMs)在支持学术评审与自动科学评估方面展现出潜力,但其精确定位错误的能力仍缺乏系统评估。本文提出FLAWS(Fault Localization Across Writing in Science),一个包含713对论文-错误组合的自动化基准,用于评估大模型识别并定位削弱核心论点错误的能力。我们通过大模型系统性地向已通过同行评审的论文中插入可推翻结论的错误,并设计自动化评估指标,衡量模型能否准确定位这些错误。该基准面临三大挑战:确保插入错误具有明确性、难度与内容相关性,避免导致识别过于简单的线索,以及实现可扩展的自动化评估。我们在该基准上评估了五种前沿大模型:Claude Sonnet 4.5、DeepSeek Reasoner v3.1、Gemini 2.5 Pro、GPT 5 和 Grok 4。结果显示,GPT 5 表现最佳,在 k=10(即模型生成前10个最可能含错的文本片段)时,识别准确率达到39.1%。
原文摘要 · Abstract (English)
The identification and localization of errors is a core task in peer review, yet the exponential growth of scientific output has made it increasingly difficult for human reviewers to reliably detect errors given the limited pool of experts. Recent advances in Large Language Models (LLMs) have sparked interest in their potential to support such evaluation tasks, from academic peer review to automated scientific assessment. However, despite the growing use of LLMs in review systems, their capabilities to pinpoint errors remain underexplored. In this work, we introduce Fault Localization Across Writing in Science (FLAWS), an automated benchmark consisting of 713 paper-error pairs designed to evaluate how effectively LLMs detect errors that undermine key claims in research papers. We construct the benchmark by systematically inserting claim-invalidating errors into peer-reviewed papers using LLMs, paired with an automated evaluation metric that measures whether models can identify and localize these errors. Developing such a benchmark presents unique challenges that we overcome: ensuring that the inserted errors are well-defined, challenging, and relevant to the content of the paper, avoiding artifacts that would make identification trivial, and designing a scalable, automated evaluation metric. On the resulting benchmark, we evaluate five frontier LLMs: Claude Sonnet 4.5, DeepSeek Reasoner v3.1, Gemini 2.5 Pro, GPT 5, and Grok 4. Among these, GPT 5 is the top-performing model, achieving 39.1% identification accuracy when k=10, where k is the number of top-ranked error text candidates generated by the LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。