通过软件共使用图发现数学软件社区并预测论文归属
Graph-Based Discovery of Mathematical Software Communities and Publication-to-Community Prediction

- 构建论文-软件共用网络,用社区检测发现数学软件群体
- 基于MSC分类的特征比标题嵌入更准确预测论文所属社区
- 适合关注科研软件发现与推荐的研究者参考
研究软件形成跨越传统学科边界的共用社区,但其结构仍缺乏探索。本文提出基于图的框架,从经整理的swMATH数据集构建软件共用网络,并应用社区检测方法揭示了数学软件社区的异质性格局。将论文-社区映射建模为多标签分类任务,进一步探究能否仅凭轻量级学术元数据预测社区归属。对比两种论文特征表示:数学学科分类(MSC)与基于标题的嵌入。在多种模型下,结构化MSC表示始终表现更优,说明结构化领域元数据比仅依赖标题语义更能有效捕捉软件社区结构。该研究凸显了结构化学术元数据在大规模科研软件发现、分类与推荐中的持续价值。
原文摘要 · Abstract (English)
Research software forms distinct co-usage communities that span traditional disciplinary boundaries, yet the structure of these communities remains largely unexplored. We present a graph-based framework for discovering mathematical software communities and predicting their association with research publications. We construct a software co-usage network from publication-software relationships using a curated swMATH dataset and subsequently apply community detection method, revealing a heterogeneous landscape of mathematical software communities. We formulate publication-to-community mapping as a multi-label classification task and further investigate whether community membership can be predicted from lightweight scholarly metadata. Specifically, we compare two feature representations of scientific publications: Mathematics Subject Classification (MSC) and title-based embeddings. Across a range of models, structured MSC representation consistently provides a stronger precision-recall trade-off, demonstrating that structured domain metadata captures software-community structure more effectively than compressed title-only semantics in this setting. This work highlights the continuing value of structured scholarly metadata for large-scale research software discovery, classification and recommendation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。