arXiv:2409.08936cs.AIcs.CL2024-09被引 7

构建1万条模拟病历,连接文本与表格数据,助力医学信息抽取研究。

SimSUM: Simulated Benchmark with Structured and Unstructured Medical Records

  • 用贝叶斯网络生成带结构化背景的模拟病历,文本由GPT-4o生成。
  • 每条记录含症状提及标注,专家评估确认文本质量可靠。
  • 适合研究多模态医疗信息抽取、临床推理自动化等任务。

临床信息抽取从非结构化医疗文本中提取概念仍具挑战性,若能结合电子病历中的结构化背景信息将获益。现有开源数据集缺乏结构化特征与文本概念之间的明确关联,亟需新数据集。我们提出SimSUM,一个包含10,000条模拟患者记录的基准数据集,将非结构化临床笔记与结构化背景变量相链接。每条记录模拟呼吸系统疾病患者的就诊过程,其表格数据(如症状、诊断、基础疾病)通过领域专家定义结构与参数的贝叶斯网络生成。大语言模型(GPT-4o)被提示生成描述该就诊经历的临床笔记,包含症状及相关上下文。这些笔记经人工标注了症状实体跨度。我们进行专家评估以验证笔记质量,并在表格和文本数据上运行基线预测模型。SimSUM主要服务于在存在结构化背景变量条件下,基于领域知识将文本中感兴趣概念(本数据集中为症状)进行抽取的研究;次要用途包括跨表格与文本的临床推理自动化、存在表格/文本混杂因素时的因果效应估计,以及多模态合成数据生成。该数据集不用于训练临床决策支持系统或生产级模型,而是为简化且可控的研究提供可复现环境。

原文摘要 · Abstract (English)

Clinical information extraction, which involves structuring clinical concepts from unstructured medical text, remains a challenging problem that could benefit from the inclusion of tabular background information available in electronic health records. Existing open-source datasets lack explicit links between structured features and clinical concepts in the text, motivating the need for a new research dataset. We introduce SimSUM, a benchmark dataset of 10,000 simulated patient records that link unstructured clinical notes with structured background variables. Each record simulates a patient encounter in the domain of respiratory diseases and includes tabular data (e.g., symptoms, diagnoses, underlying conditions) generated from a Bayesian network whose structure and parameters are defined by domain experts. A large language model (GPT-4o) is prompted to generate a clinical note describing the encounter, including symptoms and relevant context. These notes are annotated with span-level symptom mentions. We conduct an expert evaluation to assess note quality and run baseline predictive models on both the tabular and textual data. The SimSUM dataset is primarily designed to support research on clinical information extraction in the presence of tabular background variables, which can be linked through domain knowledge to concepts of interest to be extracted from the text -- namely, symptoms in the case of SimSUM. Secondary uses include research on the automation of clinical reasoning over both tabular data and text, causal effect estimation in the presence of tabular and/or textual confounders, and multi-modal synthetic data generation. SimSUM is not intended for training clinical decision support systems or production-grade models, but rather to facilitate reproducible research in a simplified and controlled setting.

医学信息抽取合成数据多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。