无需人工标注,自动提取并关联主诉中的医学实体
Weakly Supervised Medical Entity Extraction and Linking for Chief Complaints
- 用拆分匹配算法生成120万条弱标注数据
- 基于BERT的模型在无标注下实现优异实体识别与链接
- 适合医疗文本挖掘与跨机构标准化场景
主诉是患者自述的就诊原因,有助于快速了解病情,也是医学文本挖掘的重要摘要。但主诉记录因录入方式多样,导致医学术语表达差异大,难以在不同医疗机构间统一标准。本研究提出一种弱监督方法,在无人工标注条件下自动提取并链接主诉中的医学实体。首先采用拆分匹配算法,在120万条真实、去标识化且经伦理审批的主诉记录上生成弱标注(包括实体提及跨度和类别标签)。随后,利用生成的弱标签训练基于BERT的模型,定位主诉文本中的实体并将其链接至预定义本体。大量实验表明,该弱监督实体抽取与链接方法( extsc{OurS})在无任何人工标注的情况下,性能优于以往方法。
原文摘要 · Abstract (English)
A Chief complaint (CC) is the reason for the medical visit as stated in the patient's own words. It helps medical professionals to quickly understand a patient's situation, and also serves as a short summary for medical text mining. However, chief complaint records often take a variety of entering methods, resulting in a wide variation of medical notations, which makes it difficult to standardize across different medical institutions for record keeping or text mining. In this study, we propose a weakly supervised method to automatically extract and link entities in chief complaints in the absence of human annotation. We first adopt a split-and-match algorithm to produce weak annotations, including entity mention spans and class labels, on 1.2 million real-world de-identified and IRB approved chief complaint records. Then we train a BERT-based model with generated weak labels to locate entity mentions in chief complaint text and link them to a pre-defined ontology. We conducted extensive experiments, and the results showed that our Weakly Supervised Entity Extraction and Linking (\ours) method produced superior performance over previous methods without any human annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。