arXiv:2506.04045math.OCcs.LG2025-06

用模糊聚类分析海量论文,提升学术数据挖掘效率

Similarity-based fuzzy clustering scientific articles: potentials and challenges from mathematical and computational perspectives

  • 基于相似度的模糊聚类,让论文可属于多个主题组
  • 提出新优化方法,处理百万级论文与百亿级引用数据
  • 结合GPU加速,解决大规模学术数据计算瓶颈

模糊聚类允许一篇论文以软隶属度归属多个主题,对出版数据分析至关重要。该问题可建模为约束优化问题,目标是最小化数据中观察到的相似度与预测分布之间的差异。尽管该方法可借助先进优化算法,但在真实大规模数据库(如 OpenAlex、Web of Science)上应用仍面临挑战——这些数据库包含约7000万篇论文和超过10亿次引用。本文从数学与计算角度分析该方法的潜力与挑战,建立二阶最优性条件,提供新理论洞见,并通过挖掘问题结构提出实用求解方法。特别地,采用基于GPU的并行计算加速梯度投影法,有效应对大规模数据处理需求。

原文摘要 · Abstract (English)

Fuzzy clustering, which allows an article to belong to multiple clusters with soft membership degrees, plays a vital role in analyzing publication data. This problem can be formulated as a constrained optimization model, where the goal is to minimize the discrepancy between the similarity observed from data and the similarity derived from a predicted distribution. While this approach benefits from leveraging state-of-the-art optimization algorithms, tailoring them to work with real, massive databases like OpenAlex or Web of Science - containing about 70 million articles and a billion citations - poses significant challenges. We analyze potentials and challenges of the approach from both mathematical and computational perspectives. Among other things, second-order optimality conditions are established, providing new theoretical insights, and practical solution methods are proposed by exploiting the structure of the problem. Specifically, we accelerate the gradient projection method using GPU-based parallel computing to efficiently handle large-scale data.

模糊聚类学术分析大规模计算优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。