arXiv:2505.20323cs.CLcs.AI2025-05被引 2

构建了超百万条时间事件的医学病例文本时序库,助力临床轨迹研究。

PMOA-TTS: Introducing the PubMed Open Access Textual Times Series Corpus

  • 用大模型将12万份病例转为带时间戳的事件对,实现自动化结构化。
  • 包含超560万条标注事件,经临床专家验证,时间对齐误差低。
  • 适合做医疗时序分析、生存预测和自然语言理解的研究者使用。

临床叙述蕴含患者病程的时间动态,但大规模带时间标注的数据资源稀缺。我们提出PMOA-TTS,一个由124,699名患者的开放获取医学文献案例报告构建的时序语料库,通过可扩展的大语言模型管道(Llama 3.3 70B与DeepSeek-R1)转化为结构化的(事件,时间)对。该语料库包含超过560万条带时间戳的事件,并提取了人口统计学信息与诊断结果。技术验证采用临床医生人工标注的黄金数据集,结合语义事件匹配、时间一致性(c-index)及对齐误差(以对数时间累积分布函数下的曲线下面积,AULTC表示)。我们评估了不同提示策略与模型选择的影响,并提供完整文档支持复现。PMOA-TTS可推动叙事文本中的时间线抽取、时序推理、生存建模与事件预测研究,具备广泛诊断与人群覆盖。数据与代码已开源。

原文摘要 · Abstract (English)

Clinical narratives encode temporal dynamics essential for modeling patient trajectories, yet large-scale temporally annotated resources are scarce. We introduce PMOA-TTS, a corpus of 124,699 single-patient PubMed Open Access case reports converted into structured textual timelines of (event, time) pairs using a scalable large-language-model pipeline (Llama 3.3 70B and DeepSeek-R1). The corpus comprises over 5.6 million timestamped events, alongside extracted demographics and diagnoses. Technical validation uses a clinician-curated gold set and three measures: semantic event matching, temporal concordance (c-index), and alignment error summarized with Area Under the Log-Time CDF (AULTC). We benchmark alternative prompting and model choices and provide documentation to support reproduction. PMOA-TTS enables research on timeline extraction, temporal reasoning, survival modeling and event forecasting from narrative text, and offers broad diagnostic and demographic coverage. Data and code are openly available in public repositories.

医疗文本时间序列语料库大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。