arXiv:2606.11361cs.IRcs.CL2026-06

构建了2320万篇结构化医学文献摘要数据集,助力文本挖掘与知识提取。

A PubMed-Scale Dataset of Structured Biomedical Abstracts

  • 从PubMed全库提取并结构化2320万篇摘要,分两类:作者标注590万条,模型自动标注1720万条。
  • 所有摘要统一为五段式框架,保留原始文献标识、发表类型和日期信息。
  • 适合训练分类模型、测试分割算法或开展大规模领域特定信息抽取研究。

结构化摘要对生物医学文献处理至关重要,有助于信息检索、文本挖掘和知识整合。然而,大量索引于PubMed的摘要仍为非结构化状态,成为下游文本处理流程的瓶颈。为此,我们推出了Structured PubMed,一个涵盖超过2320万条研究论文记录的结构化摘要语料库,源自完整的PubMed数据库。该语料库分为两个子集:从官方XML文件解析出的590万条作者结构化摘要,以及通过逐字提取大语言模型管道自动标注的1720万条原非结构化摘要。每条记录均采用统一的五段式结构框架,并映射至其原始PubMed标识符、出版类型和出版日期。该数据集可用于训练句子分类模型、评估文本分割架构,以及在前所未有的全PubMed尺度上实现分章节信息抽取。

原文摘要 · Abstract (English)

Structured abstracts are important for biomedical literature processing, by facilitating information retrieval, text mining, and knowledge synthesis. However, a vast portion of abstracts indexed in PubMed remain unstructured, presenting a significant bottleneck for downstream text-processing workflows and applications. To resolve this limitation, we introduce Structured PubMed, a comprehensive corpus of section-labeled biomedical abstracts compiled from the complete PubMed database, encompassing over 23.2 million research-article records. The corpus is divided into two distinct subsets: a collection of 5.9 million author-structured abstracts parsed from official XML files, and an automatically labeled collection of 17.2 million originally unstructured abstracts structured via a verbatim-extraction Large Language Model pipeline. Every record is harmonized under a unified five-section schema and mapped to its original PubMed identifier, publication type, and publication date. This dataset can be utilized to train sentence-classification models, benchmark text-segmentation architectures, and perform large-scale, section-specific information extraction at an unprecedented PubMed-wide scale.

生物医学数据集结构化LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。