arXiv:2605.27072cs.CLcs.AI2026-05

E3自动检测论文中的关键问题,准确率远超人类与现有大模型。

E3: Issue-Level Backtesting for Automated Research Critique

论文配图:E3: Issue-Level Backtesting for Automated Research Critique
图 1 · 摘自论文原文
  • 基于问题级回测,自动识别论文的技术漏洞与潜在风险。
  • 在100篇ICLR 2026论文中召回率达90.2%,超越人类与主流大模型。
  • 适合审稿人、研究团队及关注论文质量的AI开发者使用。

我们提出E3,一个自动化审稿助手,能识别研究论文中影响决策的关键技术问题。对每个问题,E3报告其性质、位置、对贡献的影响,以及解决所需分析或证据,涵盖未经证实的主张、缺失消融实验、弱基线、隐藏假设、有效性威胁和数据泄露风险。为避免评估偏差,采用问题级回测协议:语料库仅包含训练截止日后发表的论文,且由匿名审稿人组成的元评审员(meta-judge)基于仅可见匿名审稿内容,将每对问题-来源标注为‘已捕获’、‘部分捕获’或‘遗漏’。在100篇ICLR 2026论文、4598个问题行上,对比E3与人工审稿及两个基于GPT-5.4和Claude-Opus-4-6的大模型基线,使用GPT-5.5作为元评审员,E3在各项聚合指标中均取得最高召回率。部分包含召回率达90.2%,较GPT高出15.5个百分点,较Claude高17.1,较人类审稿高29.2;严格召回率保持相同排序,达65.8%。对于人类审稿人提出的关切,E3覆盖89.6%;对人类遗漏的问题,额外发现1635条,比次优源多出406条。相关语料库、基线提示、评审提示模板与评估代码均已公开。

原文摘要 · Abstract (English)

We present E3, an automated review assistant that augments reviewers and engineering teams by identifying decision-relevant technical concerns in research papers. For each concern, E3 reports its nature, its location, its bearing on the contribution, and the analysis or evidence that would resolve it, covering unsupported claims, missing ablations, weak baselines, hidden assumptions, threats to validity, and leakage risks. To evaluate E3 without contamination confounds we adopt an issue-level backtesting protocol: the corpus is restricted to papers postdating the training cutoff of every automated source, and for each paper a meta-judge that observes only anonymised reviews labels every issue-source pair as Caught, Partial, or Missed. Applied to 100 ICLR 2026 papers and 4598 judged issue rows, comparing E3 against the ICLR human reviews and two prompt-matched LLM baselines built on gpt-5.4 from OpenAI and claude-opus-4-6 from Anthropic, with meta-judge gpt-5.5, E3 attains the highest recall on every aggregate metric. Partial-inclusive recall reaches 90.2 percent, which is 15.5 points over GPT, 17.1 points over Claude, and 29.2 points over the human reviews, and strict recall preserves the ordering at 65.8 percent. On concerns raised by the human reviewers, E3 recovers 89.6 percent; on concerns the human reviewers missed it surfaces 1635 additional rows admitted into the judged union, 406 above the next-best source. Corpus, baseline prompts, judge prompt template, and evaluation code are released.

自动审稿论文评估AI质检大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。