arXiv:2506.23366cs.DLcs.CL2025-06被引 1

用语义密度和不对称性预测论文影响力,发现密度有微弱但稳定信号。

Density, asymmetry and citation dynamics in scientific literature

  • 引入语义空间中的密度与不对称性度量论文与前人研究的相似性
  • 密度能小幅提升引文预测准确率,不对称性无显著作用
  • 适用于关注科学影响力建模的研究者,方法可扩展至多领域

科学研究行为常在继承已有知识与提出新思想之间存在张力。本文探究这种张力是否反映在论文与其前人研究的相似性与其后续引文率之间的关系上。为量化相似性,提出两个互补指标:(1) 密度(ρ),即固定数量先前发表论文在语义嵌入空间中所占最小距离内的比例;(2) 不对称性(α),即论文与其最近邻在方向上的平均差异。基于约5.3万篇跨九大学科、五种文档嵌入模型的论文,采用贝叶斯分层回归分析两者的预测关系。尽管ρ对引文数的单独影响较小且波动大,但加入密度特征后,模型在外部数据上的预测性能持续提升。结果表明,论文周围文献的语义密度虽仅携带微弱信号,但仍具信息价值。而出版物的不对称性未能提升引文预测能力。本研究提供了一种将文档嵌入与科学计量结果关联的可扩展框架,并引发关于语义相似性如何影响科学奖励机制的新问题。

原文摘要 · Abstract (English)

Scientific behavior is often characterized by a tension between building upon established knowledge and introducing novel ideas. Here, we investigate whether this tension is reflected in the relationship between the similarity of a scientific paper to previous research and its eventual citation rate. To operationalize similarity to previous research, we introduce two complementary metrics to characterize the local geometry of a publication's semantic neighborhood: (1) \emph{density} ($ρ$), defined as the ratio between a fixed number of previously-published papers and the minimum distance enclosing those papers in a semantic embedding space, and (2) asymmetry ($α$), defined as the average directional difference between a paper and its nearest neighbors. We tested the predictive relationship between these two metrics and its subsequent citation rate using a Bayesian hierarchical regression approach, surveying $\sim 53,000$ publications across nine academic disciplines and five different document embeddings. While the individual effects of $ρ$ on citation count are small and variable, incorporating density-based predictors consistently improves out-of-sample prediction when added to baseline models. These results suggest that the density of a paper's surrounding scientific literature may carry modest but informative signals about its eventual impact. Meanwhile, we find no evidence that publication asymmetry improves model predictions of citation rates. Our work provides a scalable framework for linking document embeddings to scientometric outcomes and highlights new questions regarding the role that semantic similarity plays in shaping the dynamics of scientific reward.

引文预测语义嵌入科学计量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。