用结构化BERT预测论文引用量,提升准确率。
CiMaTe: Citation Count Prediction Effectively Leveraging the Main Text
- 基于BERT显式建模论文章节结构,挖掘主文信息
- 在计算语言学和生物领域分别提升5.1和1.8点相关性
- 适合关注学术影响力预测的研究者
预测论文未来的引用次数在海量论文中发现有价值研究日益重要。尽管论文主文是影响引用数的关键因素,但因其内容过长,难以有效融入机器学习模型,此前研究未充分探索其利用方式。本文提出一种基于BERT的引用预测模型CiMaTe,通过显式捕捉论文的章节结构来利用主文信息。在计算语言学与生物学领域的实验表明,CiMaTe在斯皮尔曼等级相关系数上分别优于先前方法5.1分和1.8分,验证了其有效性。
原文摘要 · Abstract (English)
Prediction of the future citation counts of papers is increasingly important to find interesting papers among an ever-growing number of papers. Although a paper's main text is an important factor for citation count prediction, it is difficult to handle in machine learning models because the main text is typically very long; thus previous studies have not fully explored how to leverage it. In this paper, we propose a BERT-based citation count prediction model, called CiMaTe, that leverages the main text by explicitly capturing a paper's sectional structure. Through experiments with papers from computational linguistics and biology domains, we demonstrate the CiMaTe's effectiveness, outperforming the previous methods in Spearman's rank correlation coefficient; 5.1 points in the computational linguistics domain and 1.8 points in the biology domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。