arXiv:2603.03126cs.DLcs.DB2026-03

整合2.93亿篇论文的开源科研数据湖,实现跨源统一查询。

The Science Data Lake: A Unified Open Infrastructure Integrating 293 Million Papers Across Eight Scholarly Sources with Embedding-Based Ontology Alignment

  • 基于DuckDB和Parquet构建,通过DOI标准化统一8个开放数据源
  • 用嵌入模型对齐4516个主题到13个科学本体,映射覆盖率99.8%
  • 适合需要跨库分析的科研人员,支持本地部署与远程查询

学术数据分散在孤立数据库中,元数据不一致且关联缺失。我们提出Science Data Lake,一个基于DuckDB和Parquet文件的可本地部署基础设施,通过DOI标准化统一了8个开源来源:Semantic Scholar、OpenAlex、SciSciNet、Papers with Code、Retraction Watch、Reliance on Science、预印本到正式发表的映射关系以及Crossref。该资源包含约960GB的Parquet文件,覆盖约2.93亿篇唯一标识的论文,涉及约22种模式和约153个SQL视图。采用BGE-large句向量进行嵌入式本体对齐,将4,516个OpenAlex主题映射到13个科学本体(约130万术语),生成16,150条映射,覆盖99.8%的主题(阈值≥0.65),在推荐的≥0.85操作点上达到F1=0.77,优于TF-IDF、BM25和Jaro-Winkler基线(300对黄金标准评估)。通过10项自动化检查、跨源引用一致性分析(成对皮尔逊相关系数r=0.76–0.87)和分层人工标注验证。四个示例展示单数据库无法实现的跨源分析。资源开源,可部署于单个磁盘或通过HuggingFace远程查询,配备适合大语言模型研究代理的结构化文档。

原文摘要 · Abstract (English)

Scholarly data are largely fragmented across siloed databases with divergent metadata and missing linkages among them. We present the Science Data Lake, a locally-deployable infrastructure built on DuckDB and simple Parquet files that unifies eight open sources - Semantic Scholar, OpenAlex, SciSciNet, Papers with Code, Retraction Watch, Reliance on Science, a preprint-to-published mapping, and Crossref - via DOI normalization while preserving source-level schemas. The resource comprises approximately 960GB of Parquet files spanning ~293 million uniquely identifiable papers across ~22 schemas and ~153 SQL views. An embedding-based ontology alignment using BGE-large sentence embeddings maps 4,516 OpenAlex topics to 13 scientific ontologies (~1.3 million terms), yielding 16,150 mappings covering 99.8% of topics ($\geq 0.65$ threshold) with $F1 = 0.77$ at the recommended $\geq 0.85$ operating point, outperforming TF-IDF, BM25, and Jaro-Winkler baselines on a 300-pair gold-standard evaluation. We validate through 10 automated checks, cross-source citation agreement analysis (pairwise Pearson $r = 0.76$ - $0.87$), and stratified manual annotation. Four vignettes demonstrate cross-source analyses infeasible with any single database. The resource is open source, deployable on a single drive or queryable remotely via HuggingFace, and includes structured documentation suitable for large language model (LLM) based research agents.

科研数据知识图谱跨源集成开放数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。