arXiv:2505.08941cs.LGcs.CL2025-05被引 1

用预训练模型预测论文未来被引率,准确率达0.826

ForeCite: Adapting Pre-Trained Language Models to Predict Future Citation Rates of Academic Papers

  • 在预训练语言模型后接线性层,用于回归预测引用率
  • 在90万篇生物医学论文上达0.826相关性,提升27点
  • 适合科研评估自动化与学术影响力预测研究者

预测学术论文未来的引用次数是实现科研评价自动化和加速科学进步的重要一步。我们提出ForeCite,一个简单而强大的框架,通过在预训练因果语言模型后添加线性头来预测平均每月引用率。将Transformer应用于回归任务,ForeCite在包含90万+篇2000至2024年发表的生物医学论文的精选数据集上实现了测试相关性ρ=0.826,较此前最先进方法提升27点。全面的缩放定律分析显示,模型规模和数据量增加均带来持续提升;时间留出实验验证了实际鲁棒性。基于梯度的显著性热图表明模型可能过度依赖标题和摘要文本。这些结果确立了学术研究长期影响力的预测新基准,为科学贡献的自动化、高保真评估奠定了基础。

原文摘要 · Abstract (English)

Predicting the future citation rates of academic papers is an important step toward the automation of research evaluation and the acceleration of scientific progress. We present $\textbf{ForeCite}$, a simple but powerful framework to append pre-trained causal language models with a linear head for average monthly citation rate prediction. Adapting transformers for regression tasks, ForeCite achieves a test correlation of $ρ= 0.826$ on a curated dataset of 900K+ biomedical papers published between 2000 and 2024, a 27-point improvement over the previous state-of-the-art. Comprehensive scaling-law analysis reveals consistent gains across model sizes and data volumes, while temporal holdout experiments confirm practical robustness. Gradient-based saliency heatmaps suggest a potentially undue reliance on titles and abstract texts. These results establish a new state-of-the-art in forecasting the long-term influence of academic research and lay the groundwork for the automated, high-fidelity evaluation of scientific contributions.

引用预测语言模型科研评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。