arXiv:2605.07022cs.LG2026-05

用大模型从2250万篇论文自动构建更全面、精准的生物医学数据集

Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale

论文配图:Self-Driving Datasets: From 20 Million Papers to Nuanced Biomedical Knowledge at Scale
图 1 · 摘自论文原文
  • 基于九个生物医学本体,用大模型标注22.5万篇论文中45亿实体
  • 生成630万条结构化记录,多项数据集规模创公开纪录,错误率低于1%~7.7%
  • 支持带上下文的细粒度查询,保留实验条件等关键细节,适合药物研发

人工维护的生物医学数据库成本高、更新慢,且丢失实验背景信息,难以判断数据准确性和覆盖范围。我们证明可将PubMed自主、低成本转化为结构化数据集,其规模、精细度和准确性均优于现有数据库。提出三项核心成果:(1) 基于LLM的实体标注管道,基于九个生物医学本体,在2250万篇论文、2.5万亿词的PubMed语料中标注45亿实体,覆盖19类;(2) 混合稀疏-稠密检索系统,支持基于实体过滤的语义查询;(3) Starling多智能体深度研究系统,仅需自然语言任务描述,即可设计精准检索过滤器、推导提取模式,并输出带丰富上下文字段和原文支持段落的结构化记录。在血脑屏障渗透性、口服生物利用度、急性毒性(LD50)、基因-疾病关联、蛋白亚细胞定位及化学反应共六项任务中,生成约630万条记录(每项9.1万至300万条),部分为当前最大公开数据集。前沿模型对提取结果的拒绝率仅为0.6%-7.7%,远低于主流人工数据库错误率(如BBB_Martins为16.5%,Bioavailability_Ma为7.3%)。支持段落保留了传统表格数据库忽略的细微信息,如口服生物利用度可能受进食状态影响。整体构建了支撑人工智能驱动治疗设计的基础体系。代码与数据集:https://github.com/starling-labs/starling。

原文摘要 · Abstract (English)

Manually curated biomedical repositories -- spanning bioactivity, genomics, and chemistry -- are expensive to maintain, lag behind primary literature, and discard experimental context, obscuring nuances needed to assess data correctness and coverage. We show that PubMed itself can be autonomously and cost-effectively turned into structured datasets that are larger, more nuanced, and more accurate than the curated databases they replace. We present three coupled contributions: (1) an LLM-based entity-tagging pipeline, grounded in nine biomedical ontologies, that tags 4.5B entities across 19 categories in a 22.5M-paper, 2.5T-token PubMed corpus; (2) hybrid sparse-dense retrieval supporting entity-filtered semantic queries over the tagged corpus; and (3) Starling, a multi-agent deep research system that, given only a natural-language task description, designs precision- and recall-targeted retrieval filters, induces an extraction schema, and emits structured records with nuance-rich fields and supporting passages. Across six tasks -- blood-brain barrier permeability, oral bioavailability, acute toxicity (LD50), gene-disease associations, protein subcellular localization, and chemical reactions -- Starling produces ~6.3M records (91K-3M per task); several are, to our knowledge, the largest public datasets for their property. Frontier-model rejection of our extractions is 0.6-7.7% across tasks, far below error rates we measure on widely used curated counterparts (e.g., 16.5% on BBB_Martins, 7.3% on Bioavailability_Ma). Beyond scale and accuracy, the supporting passages carry nuance tabular databases discard -- e.g., oral bioavailability may depend on fed vs. fasted state. Together, the corpus, retrieval, and agent establish a foundation for AI-driven therapeutic design. Code and datasets: https://github.com/starling-labs/starling.

生物医学知识图谱大模型应用数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。