用6万篇论文数据,提前预测论文和期刊影响力。
Comprehensive Manuscript Assessment with Text Summarization Using 69707 articles
- 用Transformer提取标题摘要语义,融合标题与摘要信息
- 在69707篇论文上实现期刊与论文影响力的准确预测
- 适合科研评估、投稿选刊与论文改进建议生成
快速高效评估科研论文的未来影响力是作者与审稿人共同关注的问题。当前衡量学术论文影响力的主要标准是引用次数。近年来,已有诸多研究尝试预测不同引用周期内的引用数,但多数仅针对特定学科或需早期引用数据,难以用于论文发表前的早期评估。本文利用Scopus构建了一个涵盖69707篇来自99种跨学科期刊的大型综合数据集,提出一种基于深度学习的影响力分类方法,利用论文正文与元数据中的语义特征。为提炼语义特征(如标题和摘要),采用基于Transformer的语言模型进行编码,并设计文本融合层以捕捉标题与摘要间的共享信息。重点针对预出版阶段的论文,开展两类影响力预测任务:(1) 论文将发表期刊的影响力;(2) 论文自身的未来影响力。在该数据集上的大量实验表明,所提模型在影响力预测任务中表现优越,并展现出生成论文反馈与改进建议的潜力。
原文摘要 · Abstract (English)
Rapid and efficient assessment of the future impact of research articles is a significant concern for both authors and reviewers. The most common standard for measuring the impact of academic papers is the number of citations. In recent years, numerous efforts have been undertaken to predict citation counts within various citation windows. However, most of these studies focus solely on a specific academic field or require early citation counts for prediction, rendering them impractical for the early-stage evaluation of papers. In this work, we harness Scopus to curate a significantly comprehensive and large-scale dataset of information from 69707 scientific articles sourced from 99 journals spanning multiple disciplines. We propose a deep learning methodology for the impact-based classification tasks, which leverages semantic features extracted from the manuscripts and paper metadata. To summarize the semantic features, such as titles and abstracts, we employ a Transformer-based language model to encode semantic features and design a text fusion layer to capture shared information between titles and abstracts. We specifically focus on the following impact-based prediction tasks using information of scientific manuscripts in pre-publication stage: (1) The impact of journals in which the manuscripts will be published. (2) The future impact of manuscripts themselves. Extensive experiments on our datasets demonstrate the superiority of our proposed model for impact-based prediction tasks. We also demonstrate potentials in generating manuscript's feedback and improvement suggestions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。