arXiv:2605.13310cs.DLcs.DB2026-05被引 2

构建科研软件知识图谱,打通代码与论文的关联

SemRepo: A Knowledge Graph for Research Software and Its Scholarly Ecosystem

  • 将20万GitHub科研仓库元数据整合为RDF图谱
  • 链接作者、论文、数据集等信息,覆盖超8100万三元组
  • 适合研究可复现性与软件可持续性的学者使用

我们提出SemRepo,一个包含超过8100万三元组的RDF知识图谱,涵盖近20万与科研相关的GitHub仓库。该图谱记录了仓库级别元数据,如贡献者、问题和编程语言,并与外部学术知识图谱进行关联。具体而言,仓库作者链接至SemOpenAlex中的个人档案,仓库与LPWC中的学术出版物相连,研究产物(如数据集和实验)通过MLSea-KG实现关联。这种集成使跨出版物与学术成果的查询成为可能,而这些信息通常分散在不同平台。SemRepo支持现有资源难以实现的分析,包括跨仓库与出版物的溯源重建,以及系统性识别科研可复现性和软件可持续性的风险。通过将科研软件与其学术背景统一于单一图谱中,SemRepo为大规模分析科学生态系统内的软件提供了重要基础设施。

原文摘要 · Abstract (English)

We present SemRepo, an RDF knowledge graph comprising over 81 million triples describing nearly 200,000 GitHub repositories associated with scientific research. SemRepo captures repository-level metadata, such as contributors, issues, and programming languages, and interlinks this information with external scholarly knowledge graphs. In particular, repository authors are linked to their profiles in SemOpenAlex, repositories are connected to scholarly publications in LPWC, and research artifacts, such as datasets and experiments, are linked via MLSea-KG. This integration enables queries that span publications and their scholarly artifacts, which are typically fragmented across separate platforms. SemRepo supports analyses that are difficult to perform with existing resources in isolation, including provenance reconstruction across repositories and publications, as well as the systematic identification of risks to research reproducibility and software sustainability. By unifying research software with its scholarly context in a single graph, SemRepo provides an important infrastructure for large-scale analysis of software within the broader scientific research ecosystem.

知识图谱科研软件可复现性数据关联

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。