arXiv:2508.10795cs.CL2025-08Conference of the …被引 13

用AI分析论文新颖性,让审稿更客观透明

Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback

  • 分三步自动化评估论文新颖性:提取内容、检索文献、结构化对比
  • 在182篇ICLR 2025论文上,与人类判断一致率达86.5%
  • 适合需要提升审稿一致性与透明度的研究者和期刊编辑

新颖性评估是同行评审的核心环节,但在NLP等高产领域仍缺乏系统研究,且审稿人负担日益加重。本文提出一种结构化自动新颖性评估方法,通过三个阶段模拟专家审稿行为:从投稿中提取内容、检索并合成相关工作、进行结构化比较以支持证据驱动的判断。该方法基于大规模人工新颖性评审分析,捕捉了独立命题验证与上下文推理等关键模式。在182篇ICLR 2025投稿上,该方法与人类审稿人推理的一致率达到86.5%,在新颖性结论上的同意率为75.3%,显著优于现有LLM基线。生成的分析报告详尽且具备文献意识,提升了审稿一致性。结果表明,结构化LLM辅助方法可在不取代人类专家的前提下,助力更严谨、透明的同行评审。数据与代码已公开。

原文摘要 · Abstract (English)

Novelty assessment is a central yet understudied aspect of peer review, particularly in high volume fields like NLP where reviewer capacity is increasingly strained. We present a structured approach for automated novelty evaluation that models expert reviewer behavior through three stages: content extraction from submissions, retrieval and synthesis of related work, and structured comparison for evidence based assessment. Our method is informed by a large scale analysis of human written novelty reviews and captures key patterns such as independent claim verification and contextual reasoning. Evaluated on 182 ICLR 2025 submissions with human annotated reviewer novelty assessments, the approach achieves 86.5% alignment with human reasoning and 75.3% agreement on novelty conclusions - substantially outperforming existing LLM based baselines. The method produces detailed, literature aware analyses and improves consistency over ad hoc reviewer judgments. These results highlight the potential for structured LLM assisted approaches to support more rigorous and transparent peer review without displacing human expertise. Data and code are made available.

同行评审LLM应用新颖性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。