arXiv:2604.04074cs.AIcs.LG2026-04综述

用代码执行验证论文结论,提升审稿准确性

FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

  • 提取论文主张并结合文献与代码验证
  • 在354个真实主张上达84.3%准确率
  • 可减少58%审稿时间,适合科研助理使用

基于大语言模型的审稿系统通常孤立评估论文,难以验证依赖文献或代码的主张。我们提出FactReview,一个审计流程:提取与审稿相关的主张,将其与相关文献和参考检查关联,并在有代码时,在固定修复预算下执行公开的实现。在26篇互不重叠的测试论文、354个人工验证的主张上,FactReview实现了84.3%的F1得分。在相同后端、证据匹配的对比中,FactReview得分为4.72/5,比直接使用LLM审稿高出0.74分。移除执行证据会使17.0%的主张状态改变,高于移除任何其他单一证据源的影响。在审稿辅助研究中,FactReview将平均审稿时间减少58%,同时将基准主张覆盖率从87%提升至99%。FactReview支持基于证据的主张审计,接受决定仍由人类审稿人保留。代码已公开于https://github.com/DEFENSE-SEU/FactReview。

原文摘要 · Abstract (English)

Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims difficult to verify. We present FactReview, an audit pipeline that extracts review-relevant claims, grounds them in related work and reference checks, and, when code is available, executes released artifacts under a fixed repair budget. On 26 paper-disjoint test papers with 354 human-verified claims, FactReview achieves 84.3\% F1 for claim recovery. In a same-backend, evidence-matched comparison, FactReview scores 4.72/5 overall, outperforming a direct LLM reviewer by 0.74 points. Removing execution evidence changes 17.0\% of claim statuses, more than removing any other single evidence source. In a reviewer-assistance study, FactReview reduces mean review time by 58\% while increasing benchmark-claim coverage from 87\% to 99\%. FactReview supports evidence-based claim auditing, with acceptance decisions reserved for human reviewers. The code is public at https://github.com/DEFENSE-SEU/FactReview.

自动审稿代码验证证据审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。