LLM生成引用时会放大'马太效应',更倾向高被引论文。
How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- LLM生成引用时偏好高被引、新近、标题短的论文。
- 在10,000篇论文中,274,951条引用显示与真实引用语义相似度高。
- 可能加剧学术影响力不均,适合关注AI对科研影响的研究者阅读。
科学知识的传播依赖于研究人员发现并引用前人工作的方式。大型语言模型(LLMs)在科研过程中的应用为引用实践引入了新层面。然而,目前尚不清楚LLMs在多大程度上与人类引用习惯一致,其表现如何跨领域变化,以及可能如何影响引用动态。我们发现,LLMs在生成参考文献时系统性地强化了引用中的‘马太效应’,即持续偏爱已被广泛引用的论文。这一趋势在不同科学领域普遍存在,尽管各领域存在显著差异的‘存在率’(即生成引用与外部文献数据库中记录匹配的比例)。分析GPT-4o为10,000篇论文生成的274,951条引用发现,LLM推荐倾向于更近期、标题更短、作者更少的文献。尽管如此,生成引用在内容相关性上与真实引用水平相当,展现出类似的网络效应,并减少作者自我引用。这些发现表明,LLM可能通过反映并放大既有趋势,重塑引用实践,进而影响科学发现的轨迹。随着LLM日益融入科研流程,理解其在塑造科学共同体发现和继承前人工作的角色至关重要。
原文摘要 · Abstract (English)
The spread of scientific knowledge depends on how researchers discover and cite previous work. The adoption of large language models (LLMs) in the scientific research process introduces a new layer to these citation practices. However, it remains unclear to what extent LLMs align with human citation practices, how they perform across domains, and may influence citation dynamics. Here, we show that LLMs systematically reinforce the Matthew effect in citations by consistently favoring highly cited papers when generating references. This pattern persists across scientific domains despite significant field-specific variations in existence rates, which refer to the proportion of generated references that match existing records in external bibliometric databases. Analyzing 274,951 references generated by GPT-4o for 10,000 papers, we find that LLM recommendations diverge from traditional citation patterns by preferring more recent references with shorter titles and fewer authors. Emphasizing their content-level relevance, the generated references are semantically aligned with the content of each paper at levels comparable to the ground truth references and display similar network effects while reducing author self-citations. These findings illustrate how LLMs may reshape citation practices and influence the trajectory of scientific discovery by reflecting and amplifying established trends. As LLMs become more integrated into the scientific research process, it is important to understand their role in shaping how scientific communities discover and build upon prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。