arXiv:2409.16605cs.CLcs.AI2024-09被引 30

构建学术新颖性评估基准,测试大模型判断论文创新性的能力。

Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications

  • 设计跨领域学术论文对,以发布时间差判断新颖性
  • 提出RAG-Novelty方法,在6个领域上超越现有模型
  • 适合研究学术智能评估、论文评审辅助的学者

现有研究多从语义角度评估大语言模型的创造性,但学术出版物中的新颖性评估仍属空白。本文提出学术新颖性基准(SchNovel),包含来自arXiv数据集的15000组论文对,覆盖六个领域,时间跨度2至10年。每对中较晚发表的论文被视为更新颖。同时提出RAG-Novelty方法,模拟人类审稿人通过检索相似论文来评估新颖性。大量实验揭示不同LLM在新颖性判断上的表现差异,并证明RAG-Novelty优于近期基线模型。

原文摘要 · Abstract (English)

Recent studies have evaluated the creativity/novelty of large language models (LLMs) primarily from a semantic perspective, using benchmarks from cognitive science. However, accessing the novelty in scholarly publications is a largely unexplored area in evaluating LLMs. In this paper, we introduce a scholarly novelty benchmark (SchNovel) to evaluate LLMs' ability to assess novelty in scholarly papers. SchNovel consists of 15000 pairs of papers across six fields sampled from the arXiv dataset with publication dates spanning 2 to 10 years apart. In each pair, the more recently published paper is assumed to be more novel. Additionally, we propose RAG-Novelty, which simulates the review process taken by human reviewers by leveraging the retrieval of similar papers to assess novelty. Extensive experiments provide insights into the capabilities of different LLMs to assess novelty and demonstrate that RAG-Novelty outperforms recent baseline models.

大模型评估学术智能新颖性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。