arXiv:2601.19916cs.CL2026-01综述被引 2

构建论文错误检测基准,提升AI审稿的批判性与准确性

PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review

  • 设计跨章节推理错误数据集,支持长文本上下文评估
  • 集成显式错误检测与证据感知生成,使AI审稿更严格精准
  • 可训练轻量级模型,降低计算成本,适合实际部署

大型语言模型能生成流畅的同行评审意见,但在面对细微且分布式的实质性问题时,其评估往往缺乏足够的批判性。本文提出PaperAudit-Bench,包含两个部分:(1) PaperAudit-Dataset,一个涵盖单节内错误与跨节推理错误的错误数据集,专为长上下文环境下的受控评估设计;(2) PaperAudit-Review,一种融合结构化错误检测与证据感知生成的自动化评审框架,支持批判性评估。在PaperAudit-Bench上的实验表明,不同模型和检测深度间存在显著的错误可检测性差异,凸显了长上下文环境下识别此类错误的难度。相较于代表性自动化评审基线,将显式错误检测融入评审流程能产生系统性更严格、更具区分度的评估结果,验证其适用于同行评审。最后,我们证明该数据集可用于通过监督微调(SFT)和强化学习(RL)训练轻量级LLM检测器,实现低成本高效错误检测。

原文摘要 · Abstract (English)

Large language models can generate fluent peer reviews, yet their assessments often lack sufficient critical rigor when substantive issues are subtle and distributed across a paper. In this paper, we introduce PaperAudit-Bench, which consists of two components: (1) PaperAudit-Dataset, an error dataset covering both errors identifiable within individual sections and those requiring cross-section reasoning, designed for controlled evaluation under long-context settings; and (2) PaperAudit-Review, an automated review framework that integrates structured error detection with evidence-aware review generation to support critical assessment. Experiments on PaperAudit-Bench reveal large variability in error detectability across models and detection depths, highlighting the difficulty of identifying such errors under long-context settings. Relative to representative automated reviewing baselines, incorporating explicit error detection into the review workflow produces systematically stricter and more discriminative evaluations, demonstrating its suitability for peer review. Finally, we show that the dataset supports training lightweight LLM detectors via SFT and RL, enabling effective error detection at reduced computational cost.

AI审稿错误检测大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。