实证分析论文聚类算法在真实引文网络中的表现差异
Clustering scientific publications: lessons learned through experiments with a real citation network
- 用谱聚类、Louvain和Leiden算法分析70万篇论文的引文图
- 默认参数下算法效果差,需调参才能获得有意义结果
- 适合做文献计量研究或复杂网络分析的研究者参考
论文聚类可揭示文献数据库中的潜在研究结构。基于图的聚类方法如谱聚类、Louvain和Leiden算法因能有效建模引文网络而被广泛使用,但在真实数据上性能可能下降。本研究在包含约70万篇论文和460万条引文的Web of Science引文图上评估了这些算法的表现。结果显示,尽管Louvain和Leiden等可扩展算法运行高效,但其默认设置常导致较差的划分效果。在具有密集核心和松散连接论文的不均衡大规模网络中,需精细调参才能获得有意义结果。研究强调了大规模数据下方法选择与参数调优对文献计量聚类任务的重要性。
原文摘要 · Abstract (English)
Clustering scientific publications can reveal underlying research structures within bibliographic databases. Graph-based clustering methods, such as spectral, Louvain, and Leiden algorithms, are frequently utilized due to their capacity to effectively model citation networks. However, their performance may degrade when applied to real-world data. This study evaluates the performance of these clustering algorithms on a citation graph comprising approx. 700,000 papers and 4.6 million citations extracted from Web of Science. The results show that while scalable methods like Louvain and Leiden perform efficiently, their default settings often yield poor partitioning. Meaningful outcomes require careful parameter tuning, especially for large networks with uneven structures, including a dense core and loosely connected papers. These findings highlight practical lessons about the challenges of large-scale data, method selection and tuning based on specific structures of bibliometric clustering tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。