arXiv:2412.03557cs.DLcs.IR2024-12

用新指标FICE预测论文影响力,发现认知广度与引用量强相关

Freshness and Informativity Weighted Cognitive Extent and Its Correlation with Cumulative Citation Count

  • 引入寿命比和信息量加权的新指标FICE,衡量科学实体的认知范围
  • 在ACL语料库中验证,每篇论文的唯一科学实体数增长趋缓,但FICE与引用量高度相关
  • 适合关注学术影响力评估、文献计量分析的研究者使用

本文重新审视认知广度的定义——即一段文本中独特科学实体的数量。提出新颖的加权指标FICE,基于两个新权重:科学实体的生命周期比率和信息量。将每个科学实体的生命周期建模为多个高斯函数组合的时间依赖文档频率,并计算其在发表时刻t₀的累积文档频率与整个生命周期累积频率之比作为寿命比;信息量则通过标题中各科学实体文档频率的归一化得到。基于ACL语料库,我们验证了此前在其他领域观察到的现象:每篇论文中唯一科学实体数量的增长速度逐渐放缓。研究发现,FICE与该论文组的平均累计引用次数存在显著相关性。代码已公开于https://github.com/ZiheHerzWang/Freshness-and-Informativity-Weighted-Cognitive-Extent。

原文摘要 · Abstract (English)

In this paper, we revisit cognitive extent, originally defined as the number of unique phrases in a quota. We introduce Freshness and Informative Weighted Cognitive Extent (FICE), calculated based on two novel weighting factors, the lifetime ratio and informativity of scientific entities. We model the lifetime of each scientific entity as the time-dependent document frequency, which is fit by the composition of multiple Gaussian profiles. The lifetime ratio is then calculated as the cumulative document frequency at the publication time $t_0$ divided by the cumulative document frequency over its entire lifetime. The informativity is calculated by normalizing the document frequency across all scientific entities recognized in a title. Using the ACL Anthology, we verified the trend formerly observed in several other domains that the number of unique scientific entities per quota increased gradually at a slower rate. We found that FICE exhibits a strong correlation with the average cumulative citation count within a quota. Our code is available at \href{https://github.com/ZiheHerzWang/Freshness-and-Informativity-Weighted-Cognitive-Extent}{https://github.com/ZiheHerzWang/Freshness-and-Informativity-Weighted-Cognitive-Extent}

文献计量引用分析认知广度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。