用机器学习构建心脏病电子病历术语库,精准标注关键信息。
Curation of a Cardiology Interface Terminology for Highlighting Electronic Health Records using Machine Learning
- 基于SNOMED和病历数据构建初始术语库,迭代提取候选术语。
- 训练模型从病历中识别新术语,最终术语库覆盖率达74.21%。
- 适用于心脏病电子病历结构化处理,提升临床信息检索效率。
电子健康记录(EHR)包含大量复杂医学信息。为减少关键信息遗漏,本研究提出设计心脏病界面术语库(CIT),用于精准标注心脏病患者EHR中的重要内容。研究分三阶段:第一阶段构建初始CIT,整合SNOMED心血管子层级、从训练集病历中挖掘的术语及常用缩写与药物名;第二阶段通过迭代提取含初始术语的细粒度短语作为候选术语,经半自动审核后形成训练数据术语库(TCIT);第三阶段使用机器学习模型基于TCIT识别新术语,进一步扩展至最终CIT。最终在测试集上评估,覆盖率为74.21%,广度为1.68;20份随机病历的平均完整度达98.2%,平均简洁度为84.2%。
原文摘要 · Abstract (English)
Electronic health record (EHR) notes are dense medical documents containing large amounts of information, often filled with complex medical jargon. Highlighting all details in EHRs helps reduce the likelihood of missing crucial information by drawing attention to key content. This study proposes the design of a Cardiology Interface Terminology (CIT) to accurately highlight all details in EHR notes of cardiology patients. We introduce an innovative Machine Learning (ML) technique for the design of CIT. The ML technique requires training data. Manual preparation of such training data is time-consuming and expensive. The process of the CIT design includes three phases. In the first two phases, we innovatively derive a training data CIT to be used by the third phase, ML technique. We start by designing an initial CIT, composed of several components: the cardiology-related sub-hierarchies of SNOMED, other SNOMED concepts mined from EHRs of build set, and necessary components of terms e.g., medical abbreviations and medications. Utilizing an iterative process, fine-grained phrases containing initial CIT concepts are extracted from build set as CIT concept candidates. The candidate concepts are semi-automatically reviewed before being added to CIT, yielding the training data CIT, TCIT. In the third phase, a ML model is trained with TCIT to identify candidates fitting to be concepts in the CIT. This model is used to extract further concepts from build set, yielding the final CIT. The final CIT is then used to highlight the test set and evaluate the extent to which it captures details in an unseen EHR dataset. For this purpose, four evaluation metrics, coverage, breadth, completeness, and conciseness are used. The highlighted test set has a coverage of 74.21%, with a breadth of 1.68. For 20 random notes in test set, the average completeness is 98.2% and average conciseness is 84.2%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。