用大模型补全科学引文断裂网络,提升结构完整性。
Reconnecting Fragmented Citation Networks with Semantic Augmentation
- 结合引文拓扑与大模型文本相似性,补全断裂连接
- 在66万篇论文上降低碎片化率,保持学科一致性
- 适合需要可靠引文分析的研究者和指标构建者
引文图是刻画科学结构的基础工具,但常因遗漏相关文献的引用而出现断裂。本文提出一种计算高效的混合框架,融合引文拓扑与基于大语言模型(LLM)的文本相似性。利用662,369篇来自数学与运筹学/管理科学领域的Web of Science文献,通过小连通分量添加语义边,并按文本相似性重新加权现有引文。语义增强显著降低了网络碎片化程度,同时保持了学科同质性。相较于仅依赖嵌入聚类的方法,使用Leiden算法在增强图上进行聚类,在保留结构可解释性的同时实现了多尺度组织。该方法可高效扩展至大规模数据集,为强化基于引文的指标提供了实用策略,且不压缩学科边界。
原文摘要 · Abstract (English)
Citation graphs are fundamental tools for modeling scientific structure, but are often fragmented due to missing citations of scientifically connected articles. To address this issue, we propose a computationally efficient hybrid framework integrating citation topology with large language model (LLM)-based text similarity. Using 662,369 Web of Science publications in Mathematics and Operations Research & Management Science, we augment the original graph by adding semantic edges from small, disconnected components and weighting existing citations according to textual similarity. Semantic augmentation substantially reduces fragmentation while preserving disciplinary homogeneity. Compared to embedding-only clustering, cluster detection on augmented graphs using the Leiden algorithm retains structural interpretability while offering multi-scale organization. The method scales efficiently to large datasets and offers a practical strategy for strengthening citation-based indicators without collapsing disciplinary boundaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。