首个评估大模型学术新颖性判断能力的基准,助力自动审稿。
NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment

- 构建1684组论文-评审对,涵盖引言中的新颖性声明与专家评价
- 模型对新颖性判断准确率低,微调模型常犯指令遵循错误
- 提出四维评估框架,适合审稿自动化与模型评测研究者
新颖性是学术出版的核心要求,也是同行评审的重点,但投稿量激增给人工审稿带来压力。尽管大语言模型(LLMs)在生成审稿意见方面展现潜力,尤其经过同行评审数据微调的模型,但缺乏专门的评估基准限制了对其新颖性判断能力的系统性测评。为此,我们提出了NovBench,首个大规模基准,用于评估LLMs在支持人类同行评审中生成新颖性评价的能力。NovBench包含来自顶级NLP会议的1,684组论文-评审对,涵盖从论文引言中提取的新颖性描述及对应的专家撰写的新颖性评价。我们关注两个来源:引言提供标准化、明确的新颖性陈述,而专家评价则是当前人类判断的黄金标准。此外,我们提出一个四维评估框架(相关性、正确性、覆盖度、清晰度),以评估模型生成的新颖性评价质量。在通用与专用大模型上,采用不同提示策略的广泛实验表明,当前模型对科学新颖性的理解有限,且微调模型常出现指令遵循缺陷。这些发现凸显了需设计针对性微调策略,以同时提升新颖性理解与指令遵从能力。
原文摘要 · Abstract (English)
Novelty is a core requirement in academic publishing and a central focus of peer review, yet the growing volume of submissions has placed increasing pressure on human reviewers. While large language models (LLMs), including those fine-tuned on peer review data, have shown promise in generating review comments, the absence of a dedicated benchmark has limited systematic evaluation of their ability to assess research novelty. To address this gap, we introduce NovBench, the first large-scale benchmark designed to evaluate LLMs' capability to generate novelty evaluations in support of human peer review. NovBench comprises 1,684 paper-review pairs from a leading NLP conference, including novelty descriptions extracted from paper introductions and corresponding expert-written novelty evaluations. We focus on both sources because the introduction provides a standardized and explicit articulation of novelty claims, while expert-written novelty evaluations constitute one of the current gold standards of human judgment. Furthermore, we propose a four-dimensional evaluation framework (including Relevance, Correctness, Coverage, and Clarity) to assess the quality of LLM-generated novelty evaluations. Extensive experiments on both general and specialized LLMs under different prompting strategies reveal that current models exhibit limited understanding of scientific novelty, and that fine--tuned models often suffer from instruction-following deficiencies. These findings underscore the need for targeted fine-tuning strategies that jointly improve novelty comprehension and instruction adherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。