arXiv:2505.16330cs.CLcs.AI2025-05被引 15

找最优段落组合,用模型自动预测论文创新性

SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers

  • 测试不同论文段落组合,用大模型预测创新分
  • 引言+结果+讨论组合预测效果最佳
  • 适合需要自动化评审的学术研究者

创新性是学术论文的核心要素,现有方法多关注词或实体组合,难以全面评估。论文的创新内容通常分散在引言、方法、结果等不同部分。本文通过不同段落组合输入语言模型,预测创新评分,并分析最优组合。首先使用自然语言处理技术识别论文的IMRaD结构,再将不同组合(如引言+方法)输入预训练语言模型(PLM)和大语言模型(LLM),以专家评分作为真实标签。结果表明,引言+结果+讨论组合最适于评估创新性,而全文输入并无显著提升。此外,引言和结果对预测任务最为关键。

原文摘要 · Abstract (English)

Novelty is a core component of academic papers, and there are multiple perspectives on the assessment of novelty. Existing methods often focus on word or entity combinations, which provide limited insights. The content related to a paper's novelty is typically distributed across different core sections, e.g., Introduction, Methodology and Results. Therefore, exploring the optimal combination of sections for evaluating the novelty of a paper is important for advancing automated novelty assessment. In this paper, we utilize different combinations of sections from academic papers as inputs to drive language models to predict novelty scores. We then analyze the results to determine the optimal section combinations for novelty score prediction. We first employ natural language processing techniques to identify the sectional structure of academic papers, categorizing them into introduction, methods, results, and discussion (IMRaD). Subsequently, we used different combinations of these sections (e.g., introduction and methods) as inputs for pretrained language models (PLMs) and large language models (LLMs), employing novelty scores provided by human expert reviewers as ground truth labels to obtain prediction results. The results indicate that using introduction, results and discussion is most appropriate for assessing the novelty of a paper, while the use of the entire text does not yield significant results. Furthermore, based on the results of the PLMs and LLMs, the introduction and results appear to be the most important section for the task of novelty score prediction. The code and dataset for this paper can be accessed at https://github.com/njust-winchy/SC4ANM.

论文创新性语言模型段落组合自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。