arXiv:2602.06054cs.CL2026-02

分析10万篇论文评审,揭示AI研究原创性评价的真相

Are We Truly Innovating? A Qualitative and Quantitative Study of Originality in AI Research Papers

  • 基于10万份评审报告,量化分析原创性判断的实际标准
  • 发现当前LLM普遍高估新颖性且难识别改写抄袭
  • 提供可操作框架,助作者与审稿人提升评价一致性

评估人工智能研究的原创性是同行评审中最为关键却最不可靠的环节。审稿人对原创性的判断往往模糊、不一致,且依赖于不完整的前期工作对比。本文基于超过10万份来自顶级人工智能会议的同行评审报告,开展大规模数据驱动的定性与定量分析,涵盖领域快速发展的时期。通过结构化语义检索的先前工作和嵌入在专家评审意见中的信号,系统刻画了原创性在实践中被感知的方式,并识别出影响新颖性判断的关键维度。我们的分析构建了一个细粒度、基于证据的框架,为作者和审稿人提供可操作的洞察。此外,我们评估了当前大型语言模型(LLM)在原创性评估中的可靠性,发现这些模型倾向于系统性高估新颖性,且在面对改写时难以检测概念剽窃。相关数据集、训练模型与代码已公开:https://anonymous.4open.science/r/Novelty-Reviewer-365C/

原文摘要 · Abstract (English)

Assessing originality in AI research is arguably the most consequential yet least reliable step in peer review. Reviewer judgments of originality remain opaque, inconsistent, and dependent on comparisons to prior work that are often incomplete. In this paper, we present a large-scale, data-driven qualitative and quantitative analysis of research originality based on over 100,000 peer-review reports from leading AI venues, spanning a period of rapid growth in the field. Leveraging structured, semantically retrieved prior work and signals embedded in expert reviewer assessments, we systematically characterize how originality is perceived in practice and identify the key dimensions that most strongly influence novelty judgments. Our analysis yields a fine-grained, evidence-based framework that equips both authors and reviewers with actionable insights into how originality is evaluated. In addition, we evaluate the reliability of current large language model (LLM) agents in assessing originality. We find that these models tend to systematically overestimate novelty and struggle to detect conceptual plagiarism, particularly in the presence of paraphrasing. We release our dataset, trained models, and code at: https://anonymous.4open.science/r/Novelty-Reviewer-365C/.

原创性评估同行评审大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。